Scoring modes
How Publi-Score calculates the score depending on available data. Empirical data measured on 28 paired articles of the reference corpus (50 articles, neutral Gaussian distribution μ=50/σ=20).
Current methodology: v1.2 (since 8/22/2026)
What is a methodology version?
Comparison table
| ⚡ Quick | 🔬 Partial AI (abstract) | 📝 Full manual | 🤖 Full AI (PDF) | |
|---|---|---|---|---|
| Input | PMID/DOI | PMID only | PMID + PDF | PMID + PDF |
| Criteria coverage | ~53/100 | ~85/100 | 100/100 | 100/100 |
| Integrity coverage | ~80% | ~80% | 100% | 100% |
| Duration | ~2–5 sec | ~30–60 sec | ~10–30 min | ~30–60 sec |
| Objectivity | ✅ Auto | ✅ LLM | ⚠️ Human bias | ✅ LLM |
| Published in catalogue | ✅ | ✅ | ❌ | ✅ |
| Account required | No | Yes (free) | No | Yes (free) |
| Quota | Unlimited | Shared with Full AI (same quota) | Unlimited | 5/month (free) |
What the quick mode covers (and doesn't cover)
✓ What it covers
- • §2.3 Bibliometric impact — 100% (citations, h-index)
- • §2.6 Freshness — ~69% (publication date)
- • §2.1 Level of evidence — ~57% (study type, randomisation)
- • Retractions — 100% (PubMed API + Retraction Watch)
- • Alert signals — 100% (predatory journals, EoC)
Total: ~53% of criteria · ~80% integrity
⚠ What quick mode doesn't cover
- • Real ITT analysis (intention to treat)
- • Compliance with pre-registered protocol
- • Raw data sharing
- • Clinical benefit/risk ratio
- • §2.7 Reporting quality — 0% (requires PDF)
~47% of criteria not evaluable without PDF
The Partial AI mode without PDF
Partial AI (abstract) mode uses only the abstract and article metadata — without PDF. AI evaluates 24 subcriteria out of 26 (92% coverage, 2 remaining subcriteria not evaluable: raw data sharing and code sharing). Fidelity 60% vs full mode (65% concordance × 24/26 subcriteria, measured on 28 paired articles of the reference corpus).
60%
Fidelity (concordance × coverage)
24/26
Evaluable sub-criteria (92%)
65%
Concordance with PDF (tier avg.)
2/26
criterion not evaluable without PDF
Why 60% and not 100%?
- • The abstract doesn't contain everything — some elements of the PDF remain inaccessible (figures, tables, detailed methods, references). AI evaluates what it can document from the abstract, but stays conservative when information is missing.
- • §2.4 Reproducibility & transparency — the « raw data sharing » and « code sharing » subcriteria require examining the PDF (Data Availability sections) or associated repositories. Inaccessible from abstract alone: always 0 pt.
This mode produces a detailed score with per-criterion justifications — it is partial, not degraded. Fidelity = Concordance × Coverage = 65% × 24/26 = 60%. Close tier in 86% of cases. Activates automatically when the PDF is not open access.
Activation: automatic fallback if OA PDF unavailable. Measured on a representative corpus of 50 thematic articles (COVID, Mental Health, Nutrition, Cardio, Onco, Infectious, MetaEpistemo, Genetics — balanced A→E tier distribution reflecting ordinary medical science).
Quick algorithmic mode
Quick mode uses only public metadata (PubMed, OpenAlex, Semantic Scholar…) and algorithmic rules — no AI. It evaluates the 13 subcriteria accessible via public databases out of 26 (50% coverage). Fidelity 40% vs full mode (76% concordance × 13/26 subcriteria, measured on 28 paired articles of the reference corpus).
40%
Fidelity (concordance × coverage)
13/26
Evaluable sub-criteria (50%)
76%
Concordance with PDF (tier avg.)
13/26
PDF-only subcriteria not evaluable
Why 40% Fidelity?
Fidelity = Concordance × Coverage. Quick mode reaches 76% concordance with the full mode (PDF) on the n=28 paired corpus (close tier in 100% of cases, exact tier in 52%) — excellent stability thanks to quantitative signals from public databases (citations, journal, h-index). But it only covers 13/26 subcriteria (50% coverage). Fidelity = 76% × 50% = 38% → rounded 40%.
Quick mode remains useful for immediate orientation (close tier in 100% of cases, exact tier in 52%). But for clinical decision, full mode (PDF) remains the reference.
How we measure Fidelity
Fidelity = Concordance × Coverage. This formula measures how closely a mode reproduces the reference scoring (full PDF), accounting for both precision (concordance) and the analyzed scope (coverage).
Breakdown
- Concordance = average (strict tier + tier ±1) — % agreement with the reference PDF scoring
- Coverage = number of evaluable sub-criteria / 26 (full grid)
- Fidelity = product of both, expressed as %
Detailed calculations (3 modes)
- Full PDF (AI) — reference: 100% × 26/26 = 100%
- Abstract (AI): 65% × 24/26 = 60%
- Metadata (Quick): 76% × 13/26 = 40%
Concordance details
- Strict tier: % of articles with the same letter (A/B/C/D/E)
- Tier ±1: % of articles within ≤1 tier (e.g. B vs C, or C vs D — close but not exact)
- Concordance = (strict_tier + tier_±1) / 2
Measurement corpus (representative distribution)
Measurement performed on a balanced 50-article thematic corpus, reflecting the natural distribution of ordinary medical science:
- ~10% excellent articles (tier A — rare studies with full transparency, pre-registration, open data)
- ~28% good articles (tier B — solid design, manageable COI)
- ~18% mediocre articles (tier C — decision-making tier)
- ~29% weak articles (tier D — significant methodological limits)
- ~10% very weak articles (tier E — barely usable)
This distribution comes from a Gaussian model centered on 50/100, reflecting ordinary medical science after peer-review filtering. Measuring fidelity on an unbalanced corpus would give an unrepresentative estimate of real-world use.
Empirical data
Measured on a representative corpus of 50 thematic articles (COVID, Mental Health, Nutrition, Cardio, Onco, Infectious, MetaEpistemo, Genetics — balanced A→E tier distribution reflecting ordinary medical science).
4 critical overestimation cases in quick mode
These 4 articles are among the most viewed in the corpus. Quick mode assigns them tier A or B, while Full AI mode reveals tier D.
| Article | ⚡ Quick | 🔬 Partial AI (abstract) | 🤖 Full AI (PDF) | Gap Q→F | Main reason |
|---|---|---|---|---|---|
| Polack/Pfizer — NEJM 2020 | A | E | D | −51 pts | Major industrial COI + short editorial delay not captured |
| Voysey/AZ — Lancet 2021 | B | D | D | −38 pts | AstraZeneca COI + adaptive design + data not shared |
| Hammond/Paxlovid — NEJM 2022 | B | E | D | −37 pts | Industry-only trial + raw data unavailable |
| Molnupiravir — NEJM 2022 | B | D | D | −33 pts | Merck/Ridgeback trial — non-public data |
Why quick mode overestimates: it relies only on criteria accessible via public databases. Criterion §2.3 (bibliometric impact) weighs about 10 points and is maximal for articles published in NEJM/Lancet — which are often industry-funded trials (Pfizer, AZ, Merck) whose significant conflicts of interest are only visible upon in-depth analysis of the full text.
Which mode to choose?
| Context | Recommended mode |
|---|---|
| First exploration, monitoring | ⚡ Quick |
| Clinical decision, citation, teaching | 🤖 Full AI (PDF) |
| PDF unavailable (NEJM, Lancet…) | 🔬 Partial AI (abstract) |
| Personal learning, methodological exploration | 📝 Full manual |
