Methodology
Generated from config/ranking.json and the live corpus. Change the config, re-run the pipeline, and this page moves with the rankings — it cannot describe a formula the site is not using.
Why not just sort by score?
Community averages are not comparable across titles. A 9.1 from 400 people and a 9.1 from 400,000 are different claims about the world, and a raw average treats them identically. The two common alternatives are both broken too: sorting by popularity measures marketing budget, and “top rated” lists reward small, self-selecting audiences. Three steps address that — shrink, blend, then cut on percentile.
1 · Bayesian shrink
Each title's mean is pulled toward the corpus mean in proportion to how little evidence supports it.
bayes = (v / (v + m)) · R + (m / (v + m)) · C
| R | Title mean, recomputed from the full score histogram | per title |
|---|---|---|
| v | Users who actually scored it | per title |
| m | Prior strength — the vote count at which a title is trusted half as much as the prior | 1,500 |
| C | Vote-weighted corpus mean | 75.97 |
A title with exactly m ratings lands halfway between its own average and the corpus average. At 10·m the prior contributes under 10%.
2 · Blend five orthogonal signals
Each is z-scored over the eligible cohort, clamped to ±4σ, and combined by weight.
| Signal | Weight | What it captures | |
|---|---|---|---|
bayesScore | 0.55 | Vote-count-shrunk community rating. The backbone. | |
loveIntensity | 0.16 | favourites / popularity. Separates 'people rated it well' from 'people love it'. | |
reach | 0.12 | log10(popularity) percentile. A masterpiece nobody finished is not the same object as one that moved a generation. | |
completion | 0.10 | completed / (completed + dropped + paused). Punishes shows people abandon. | |
consensus | 0.07 | Inverse of score-distribution spread. Rewards broad agreement; polarising titles are flagged, not buried. |
These are deliberately not redundant. bayesScore says people rated it well; loveIntensitysays people cared; reach says it mattered to a lot of people; completion catches shows that are pleasant for three episodes and a chore for twelve; consensus separates “broadly good” from “loved and hated in equal measure”.
2b · Normalise signals that drifted with the instrument
AniList's user base grew by orders of magnitude between 2000 and today, so raw popularity measures whena title aired at least as much as how far it travelled. Signals listed under normalizeWithin are z-scored inside their cohort instead of across the corpus.
| Signal | Normalised within | Cohorts |
|---|---|---|
reach | year | 27 |
3 · Cut tiers on percentile
| Tier | Name | Top % | Titles | Meaning |
|---|---|---|---|---|
| S+ | Masterpiece | 0.5% | 38 | Medium-defining. The titles a decade is remembered by. |
| S | Elite | 2.5% | 148 | Exceptional on every axis; essential viewing within its genre. |
| A | Excellent | 10% | 555 | Strongly recommended without caveats. |
| B | Great | 25% | 1,112 | Very good; worth your time if the premise appeals. |
| C | Good | 50% | 1,852 | Solid, competent, unremarkable. Genre fans only. |
| D | Mediocre | 80% | 2,222 | Notable flaws; watch only with a specific reason. |
| E | Weak | 100% | 1,482 | Poorly received across the board. |
Cuts are inclusive at the boundary: a title at exactly the 50th percentile is the last C, not the first D.
Cohort tiers
A 2003 OVA judged against Fullmetal Alchemist: Brotherhood tells you nothing, so the same percentile cut is applied within each of year</code>, <code>era</code>, <code>season</code>, <code>seasonKey</code>, <code>format</code>, <code>genre</code>, <code>studio</code>, <code>lengthClass</code>, <code>origin</code>, <code>source</code>, <code>demographic. A cohort smaller than 25 titles gets a rank but no tier, because a tier over eight items is noise. This is why a title can be B globally and S within its season.
Score inflation over time
Even after cohort-normalising reach, later years hold more of the top tiers. Part is real — budgets and adaptation pipelines improved — and part is measurement: community scores drift upward everywhere, and a show that finished six months ago has not had its honeymoon scores decay. Here is the raw, uncorrected evidence.
| Year | Ranked | Median raw score | Median ratings | |
|---|---|---|---|---|
| 2000 | 92 | 6.76 | 1,651 | |
| 2001 | 127 | 6.79 | 1,331 | |
| 2002 | 142 | 6.64 | 1,588 | |
| 2003 | 148 | 6.58 | 1,649 | |
| 2004 | 178 | 6.63 | 1,533 | |
| 2005 | 191 | 6.77 | 1,845 | |
| 2006 | 232 | 6.69 | 1,618 | |
| 2007 | 234 | 6.72 | 2,047 | |
| 2008 | 231 | 6.86 | 2,988 | |
| 2009 | 261 | 6.82 | 3,024 | |
| 2010 | 253 | 6.77 | 4,015 | |
| 2011 | 305 | 6.78 | 3,685 | |
| 2012 | 350 | 6.80 | 3,686 | |
| 2013 | 330 | 6.80 | 5,336 | |
| 2014 | 389 | 6.80 | 4,978 | |
| 2015 | 347 | 6.85 | 6,135 | |
| 2016 | 418 | 6.69 | 4,470 | |
| 2017 | 376 | 6.83 | 4,279 | |
| 2018 | 383 | 6.90 | 5,022 | |
| 2019 | 339 | 6.88 | 4,712 | |
| 2020 | 298 | 6.99 | 4,235 | |
| 2021 | 356 | 7.17 | 4,477 | |
| 2022 | 318 | 7.21 | 4,494 | |
| 2023 | 327 | 7.27 | 7,545 | |
| 2024 | 297 | 7.26 | 6,933 | |
| 2025 | 284 | 7.13 | 5,428 | |
| 2026 | 203 | 7.30 | 3,160 |
Use the cohort tier for cross-era comparison. A title's year and season cohort tier is computed against its contemporaries and is immune to this drift; the global tier is not.
Eligibility
| Rule | Setting | Why |
|---|---|---|
| Minimum ratings | 400 | Below this a mean is noise. Guards against five-vote perfect scores. |
| Release status | FINISHED or RELEASING | Unaired titles have no reception to measure. |
| Adult content | excluded | Kept in the dataset and flagged; never shown here. |
| Formats | MUSIC excluded | Music videos are not comparable to narrative works. |
Exclusions are recorded on the record, never silently dropped — every excluded title carries a stated reason on its own page.
Known limitations
- Single-source reception. Scores come from AniList, whose user base skews younger, more online and more Western than MyAnimeList's, and far more than a Japanese domestic audience. Titles beloved in Japan and unlicensed abroad are systematically under-measured.
- Recency is provisional. Scores for currently-airing shows are unstable and typically inflated; they settle downward for months after a finale.
- Reach is not merit. It carries 12% deliberately — enough to keep a beloved obscurity from outranking a generation-defining work on a technicality, not enough to make this a popularity chart.
- Sequels inherit goodwill. A season 3 is rated by people who already liked seasons 1–2, so franchises skew high. The one-per-franchise view removes that effect.
- Tags are crowd-sourced and noisier on obscure titles.