Grille v1.2

Scoring modes

How Publi-Score calculates the score depending on available data. Empirical data measured on 28 paired articles of the reference corpus (50 articles, neutral Gaussian distribution μ=50/σ=20).

Current methodology: v1.2 (since 8/22/2026)

What is a methodology version?

Comparison table

⚡ Quick🔬 Partial AI (abstract)📝 Full manual🤖 Full AI (PDF)
InputPMID/DOIPMID onlyPMID + PDFPMID + PDF
Criteria coverage~53/100~85/100100/100100/100
Integrity coverage~80%~80%100%100%
Duration~2–5 sec~30–60 sec~10–30 min~30–60 sec
Objectivity✅ Auto✅ LLM⚠️ Human bias✅ LLM
Published in catalogue
Account requiredNoYes (free)NoYes (free)
QuotaUnlimitedShared with Full AI (same quota)Unlimited5/month (free)

What the quick mode covers (and doesn't cover)

What it covers

  • §2.3 Bibliometric impact — 100% (citations, h-index)
  • §2.6 Freshness — ~69% (publication date)
  • §2.1 Level of evidence — ~57% (study type, randomisation)
  • Retractions — 100% (PubMed API + Retraction Watch)
  • Alert signals — 100% (predatory journals, EoC)

Total: ~53% of criteria · ~80% integrity

What quick mode doesn't cover

  • Real ITT analysis (intention to treat)
  • Compliance with pre-registered protocol
  • Raw data sharing
  • Clinical benefit/risk ratio
  • §2.7 Reporting quality — 0% (requires PDF)

~47% of criteria not evaluable without PDF

The Partial AI mode without PDF

Partial AI (abstract) mode uses only the abstract and article metadata — without PDF. AI evaluates 24 subcriteria out of 26 (92% coverage, 2 remaining subcriteria not evaluable: raw data sharing and code sharing). Fidelity 60% vs full mode (65% concordance × 24/26 subcriteria, measured on 28 paired articles of the reference corpus).

60%

Fidelity (concordance × coverage)

24/26

Evaluable sub-criteria (92%)

65%

Concordance with PDF (tier avg.)

2/26

criterion not evaluable without PDF

Why 60% and not 100%?

  • The abstract doesn't contain everything — some elements of the PDF remain inaccessible (figures, tables, detailed methods, references). AI evaluates what it can document from the abstract, but stays conservative when information is missing.
  • §2.4 Reproducibility & transparency — the « raw data sharing » and « code sharing » subcriteria require examining the PDF (Data Availability sections) or associated repositories. Inaccessible from abstract alone: always 0 pt.

This mode produces a detailed score with per-criterion justifications — it is partial, not degraded. Fidelity = Concordance × Coverage = 65% × 24/26 = 60%. Close tier in 86% of cases. Activates automatically when the PDF is not open access.

Activation: automatic fallback if OA PDF unavailable. Measured on a representative corpus of 50 thematic articles (COVID, Mental Health, Nutrition, Cardio, Onco, Infectious, MetaEpistemo, Genetics — balanced A→E tier distribution reflecting ordinary medical science).

Quick algorithmic mode

Quick mode uses only public metadata (PubMed, OpenAlex, Semantic Scholar…) and algorithmic rules — no AI. It evaluates the 13 subcriteria accessible via public databases out of 26 (50% coverage). Fidelity 40% vs full mode (76% concordance × 13/26 subcriteria, measured on 28 paired articles of the reference corpus).

40%

Fidelity (concordance × coverage)

13/26

Evaluable sub-criteria (50%)

76%

Concordance with PDF (tier avg.)

13/26

PDF-only subcriteria not evaluable

Why 40% Fidelity?

Fidelity = Concordance × Coverage. Quick mode reaches 76% concordance with the full mode (PDF) on the n=28 paired corpus (close tier in 100% of cases, exact tier in 52%) — excellent stability thanks to quantitative signals from public databases (citations, journal, h-index). But it only covers 13/26 subcriteria (50% coverage). Fidelity = 76% × 50% = 38% → rounded 40%.

Quick mode remains useful for immediate orientation (close tier in 100% of cases, exact tier in 52%). But for clinical decision, full mode (PDF) remains the reference.

How we measure Fidelity

Fidelity = Concordance × Coverage. This formula measures how closely a mode reproduces the reference scoring (full PDF), accounting for both precision (concordance) and the analyzed scope (coverage).

Breakdown

  • Concordance = average (strict tier + tier ±1) — % agreement with the reference PDF scoring
  • Coverage = number of evaluable sub-criteria / 26 (full grid)
  • Fidelity = product of both, expressed as %

Detailed calculations (3 modes)

  • Full PDF (AI) — reference: 100% × 26/26 = 100%
  • Abstract (AI): 65% × 24/26 = 60%
  • Metadata (Quick): 76% × 13/26 = 40%

Concordance details

  • Strict tier: % of articles with the same letter (A/B/C/D/E)
  • Tier ±1: % of articles within ≤1 tier (e.g. B vs C, or C vs D — close but not exact)
  • Concordance = (strict_tier + tier_±1) / 2

Measurement corpus (representative distribution)

Measurement performed on a balanced 50-article thematic corpus, reflecting the natural distribution of ordinary medical science:

  • ~10% excellent articles (tier A — rare studies with full transparency, pre-registration, open data)
  • ~28% good articles (tier B — solid design, manageable COI)
  • ~18% mediocre articles (tier C — decision-making tier)
  • ~29% weak articles (tier D — significant methodological limits)
  • ~10% very weak articles (tier E — barely usable)

This distribution comes from a Gaussian model centered on 50/100, reflecting ordinary medical science after peer-review filtering. Measuring fidelity on an unbalanced corpus would give an unrepresentative estimate of real-world use.

Empirical data

Measured on a representative corpus of 50 thematic articles (COVID, Mental Health, Nutrition, Cardio, Onco, Infectious, MetaEpistemo, Genetics — balanced A→E tier distribution reflecting ordinary medical science).

4 critical overestimation cases in quick mode

These 4 articles are among the most viewed in the corpus. Quick mode assigns them tier A or B, while Full AI mode reveals tier D.

Article⚡ Quick🔬 Partial AI (abstract)🤖 Full AI (PDF)Gap Q→FMain reason
Polack/Pfizer — NEJM 2020AED−51 ptsMajor industrial COI + short editorial delay not captured
Voysey/AZ — Lancet 2021BDD−38 ptsAstraZeneca COI + adaptive design + data not shared
Hammond/Paxlovid — NEJM 2022BED−37 ptsIndustry-only trial + raw data unavailable
Molnupiravir — NEJM 2022BDD−33 ptsMerck/Ridgeback trial — non-public data

Why quick mode overestimates: it relies only on criteria accessible via public databases. Criterion §2.3 (bibliometric impact) weighs about 10 points and is maximal for articles published in NEJM/Lancet — which are often industry-funded trials (Pfizer, AZ, Merck) whose significant conflicts of interest are only visible upon in-depth analysis of the full text.

Which mode to choose?

ContextRecommended mode
First exploration, monitoring⚡ Quick
Clinical decision, citation, teaching🤖 Full AI (PDF)
PDF unavailable (NEJM, Lancet…)🔬 Partial AI (abstract)
Personal learning, methodological exploration📝 Full manual