Analysis

The specification the ratings were produced under, in full: the rating scale, the panel, the rounds, the aggregation rule, what may and may not be published from numbers of this reliability, and the problems known to be outstanding. It is rendered from analysis/README.md in this repository, the same document the pipeline is run against.
Nations × capabilities
197 × 288
56,736 rated cells
Coders
5
four vendors, three jurisdictions, two rounds
Measured cost
$104.39
$0.0018 per rated cell
Built
2026-09-19
candidate ratings, not adopted ratings

The tier on the home page, and the Pareto analysis below

The home table carries a Tier column. It is a Pareto domination tier over the nine domain scores: one nation dominates another when it is at least as good on all nine and better on at least one, within a quarter of a scale point. Strip the nations nobody dominates and you have tier 1; strip them and repeat for tier 2, and so on.

Section 7 below runs Pareto domination too, but never on that vector. It sorts on a domain's capability-group scores for each of the nine domain events, on the three domain scores inside a dimension, and on the three dimension scores for its overall tier. The home page's tier is a tenth cut, computed by this site, and the two do not agree — which is the point of saying so here rather than letting the word “Pareto” imply they are the same number.

Sorted onCriteriaTier 1TiersWhere it appears
The nine domain scores91123The home table's Tier column
The three dimension scores3160§7, the overall tier

The difference is not a disagreement about the data; it follows from the arithmetic. Adding criteria makes domination harder, because a nation now has to win on all nine rather than on three averages that can hide a weak domain. So the nine-domain frontier is wide, at 11 nations rather than one, and there are 23 tiers beneath it rather than 60. The nine-domain tier is the stricter reading of a nation's shape; the three-dimension tier is the sharper discriminator, and §7 is right that neither is a league table.

The same non-dominated sort is used either way, andscripts/build-ratings.py proves it on every build: it re-runs the sort on the three dimension scores and stops the build unless all 197 nations land in exactly the tier analysis/outputs/ginc_pareto_tier_ranking_2026.csv published. The tolerance is §7's own 0.25 scale points, not a figure this site chose.

Tier 1 over the nine domains — 11 nations

🇫🇮Finland🇫🇷France🇩🇪Germany🇯🇵Japan🇳🇱Netherlands🇸🇬Singapore🇰🇷South Korea🇸🇪Sweden🇨🇭Switzerland🇬🇧United Kingdom🇺🇸United States

Alphabetically. Nothing separates them: a tier is a set of nations none of which dominates another, so it carries no internal order. Where the home table sorts by Tier it orders the index inside a tier only to give the rows somewhere to sit.

How the 197 nations fall across the 23 tiers

T111
T29
T37
T410
T513
T611
T710
T813
T913
T109
T1111
T126
T1310
T149
T1510
T1610
T177
T188
T194
T204
T216
T225
T231

GINC National Capability Index 2026

A complete, reproducible rating of 197 nations against 288 capabilities across the three GINC dimensions, produced by a panel of five independent language-model coders from four vendors and three jurisdictions, in two rounds, for a measured $104.39.

Live dashboard: https://claude.ai/artifact/1Zp4dpyYZjiUHiTgrokGWY

These are candidate ratings, not adopted ratings. They are simulated expert codings produced under the GINC protocol by language models, with no human adjudication and no source documents attached to any cell. Round-1 spread, not round-3 agreement, is the honest reliability estimate.


1. What exists

Capabilities 288 (Hard Power 148, Soft Power 71, Economic Power 69)
Nations 197 (the GINC master list, ISO3 as the join key)
Rated cells 56,736
Coder judgements ~397,000
Coders 5, across 4 vendors and 3 jurisdictions
Rounds 2 (independent coding, then panel adjudication)
Measured cost $104.39

The register

Dimension Domain Groups Capabilities
Hard Power Critical Technology 10 47
Hard Power Strategic Infrastructure 10 62
Hard Power National Security 10 39
Soft Power Human Capital 9 25
Soft Power Information & Influence 6 18
Soft Power Governance & Integrity 8 28
Economic Power Financial Strength 6 24
Economic Power Production & Innovation 7 23
Economic Power Trade & Investment 7 22

Built from the v2 framework JSON in frameworks/. The Factbase v1 tree was discarded: several of its entries were indicators rather than capabilities (R&D Intensity, GDP Scale & Growth) and several were normatively contested in a way that collides with the protocol’s sixth principle.


2. The rating scale

Seven bands of three on a 0–20 range.

Band Range Test
Frontier 18–20 Does the world copy this?
Advanced 15–17 Among the world’s strongest, holds under stress
Established 12–14 Works at national scale, institutionalised, reliable
Transitional 9–11 Works unevenly; strong in the capital, thin beyond
Foundational 6–8 Basic machinery stands but narrow and fragile
Nascent 3–5 Pilots and first units only
Inception 0–2 Absent 0, Declared 1, Initiated 2

Coders climb from the bottom and stop at the first no; they judge the band first and place the integer within it second. Criterion-referenced throughout: no regional curve, no home-country lift, no halo from national reputation, no credit across capability seams, no size credit, no regime judgement, no recency inflation.

Every capability carries an explicit seam instruction naming what belongs to the neighbouring domain. Critical Technology rates the capacity to originate, design and manufacture; Strategic Infrastructure rates the deployed estate and the capacity to run it; National Security rates organised force. A nation that buys a whole system from abroad and runs it well scores high on infrastructure and low on technology.


3. The panel

Coder Model Vendor Jurisdiction Rounds Cost
coder2 deepseek-chat (Flash) DeepSeek China 1, 3 $1.28
coder3 mistral-large-latest Mistral AI France 1, 3 $2.83
coder4 claude-opus-5 Anthropic United States 1 $74.26
coder5 deepseek-v4-pro DeepSeek China 1 $22.10
coder6 gpt-5.6-luna OpenAI United States 1, 3 $3.91

Four adjudicate — DeepSeek Flash, Mistral Large, DeepSeek V4 Pro and GPT-5.6 Luna. Claude Opus 5 has no round-3 pass over the full register and codes round 1 only; its judgement is shown to every adjudicator and its bias stays measurable, but it does not decide. Opus carried 82–90 per cent of the bill on the Critical Technology run and, on the bias and convergence evidence, contributed no more signal than the cheaper models.

Panel parity carries no consequence here because the adjudicated score is a rounded mean rather than a median (see §5). Under a median it would: with four adjudicators the two middle scores differ on 20 per cent of cells, and any tie rule shifts the whole index by about 0.105 points, roughly one coder standard deviation.

Choosing the OpenAI seat

Four OpenAI models were tested on two capabilities × 197 nations before one was adopted.

Model $ / 288 caps, both rounds MAE vs panel Bias Floor use
gpt-5.6-luna $3.90 0.71 / 0.84 −0.05 / −0.46 panel-normal
gpt-5.6-terra $29.33 0.82 / 0.85 +0.23 / +0.08 slightly high
gpt-5.6-sol $67.83 1.27 / 1.07 +1.17 / +0.90 poor
gpt-6-astra $105.20 1.59 / 1.61 +1.52 / +1.49 fails

Agreement and protocol discipline get worse as price rises. On judicial independence the panel places 10 nations in Inception and 39 in Nascent; gpt-6-astra places zero and eight. It refuses to use the bottom third of the scale while being marginally harsher at the top — pure scale compression, which is the one thing an index built for discrimination cannot tolerate. Raw test data in outputs/openai_coder_test_2026.csv.


4. Rounds

Round 1 — independent. Each coder scores all 197 nations on a capability without seeing any other coder. Batched 25–50 nations per call.

Round 3 — panel adjudication. Each adjudicator is shown the full panel’s round-1 table for that capability, with region and sub-region, and asked for its own adjudicated score. The prompt instructs it to adjudicate on evidence rather than average mechanically, to check intra-region ordering and cross-region comparability, and explicitly that “convergence is not the goal; a defensible outlier is worth more than a false consensus.”

Round 2 was dropped. On the Energy Technologies pilot, self-refinement moved the panel almost not at all: mean cross-coder SD went 0.98 → 0.98, with essentially all movement arriving in round 3 when coders could see each other.


5. Aggregation

The published capability score is an integer

The adjudicated score for a nation on one capability is the mean of the four adjudicators’ round-3 integers, rounded to the nearest integer — one integer out of 20, exactly as a human panel would record it.

The rule was tested both ways and depends on panel size:

Adjudicators Rule Cells carrying a score no adjudicator gave
3 median 0 per cent
3 rounded mean 4.08 per cent
4 rounded mean 1.95 per cent
4 median needs a tie rule; the two middle scores differ on 20 per cent of cells

With three seats the median wins: a rounded mean returned a value nobody had chosen on one cell in twenty-five, including cases like a panel of 4, 5 and 14 averaging to 8 — a band two of the three rejected. With four seats the mean wins: it uses every seat rather than discarding the extremes, it needs no parity tie-break, and the “nobody chose this” rate more than halves.

The cost of the mean is a small loss of outlier robustness on the 2.7 per cent of cells where the panel spreads three points or more. That is why the panel standard deviation travels with every published figure.

Equal weight at every level

capability → capability group → domain → dimension → index

An unweighted mean at each step. Hard, Soft and Economic Power each count for a third; within each dimension every domain counts once; within each domain every group counts once.

A flat mean over all 288 capabilities would instead hand Hard Power 51.4 per cent of the index, Soft Power 24.7 and Economic Power 24.0, purely because Hard Power happens to be cut into more capabilities. Nobody chose that. The fix barely moves the leading group and corrects the tail: North Korea falls 17 places once its hard-power-heavy profile stops being over-counted.


6. What to publish: bands, not ranks

Rebuilding the whole index separately from each adjudicator’s own scores gives a median standard deviation of 0.124. Against that:

Granularity Nations within 1 SD of a boundary
7 protocol bands 20 of 197 (90 per cent unambiguous)
21 steps with +/− modifiers 45 of 197
Ranks at ranks 31–60 the median gap between adjacent nations is 0.073 against a coder SD of ~0.19

Publish the band. Order alphabetically within it. Never sort by score inside a band — the moment you do, readers quote the order you just disclaimed. Carry ±SD and a boundary flag on every number.

The 2026 bands

Band n Nations
Frontier 1 United States (18.03)
Advanced 10 Canada, China, France, Germany, Japan, Netherlands, Singapore, South Korea, Sweden, United Kingdom
Established 17
Transitional 28
Foundational 40
Nascent 74
Inception 27

The United States sits at 18.03, barely over the Frontier boundary and flagged as within one standard deviation of it. Under the median rule it scored 17.90 and the Frontier band was empty. Treat “the United States is the only Frontier nation” as a statement the method cannot resolve.


7. The second ranking: Pareto tiers

An alternative that needs no weighting at all.

Within each of the nine domains a nation is described by its vector of capability-group scores. Nation A dominates B when A is at least as good on every group and better on at least one. Peeling off the non-dominated set repeatedly gives tiers: tier 1 is the frontier, tier 2 what is left once the frontier is removed, and so on. Nations are ranked lexicographically on how many domains they hold at tier 1, then tier 2, then tier 3, and down through every tier — which separates all 197 with no ties. A tolerance of 0.25 scale points applies, without which rounding noise puts dozens of nations in tier 1 everywhere.

The same comparison is run inside each dimension (three domains each) and inside each domain (tier, then nations dominated).

Use it as a diagnostic, not a league table. It cannot see level: North Korea rises a long way on a single mid-tier domain against an overall index near 4. It is fragile to noise where the mean absorbs it. And it depends on a tolerance parameter that moves the table.

What it is genuinely good for: identifying who holds the frontier (17 of 197 nations hold any of the 41 tier-1 seats across nine domains) and reading a nation’s shape rather than its size. Among the top thirty, the correlation between how far a nation’s strongest domain sits above its own mean and how far it rises on the tier rule is +0.58. Rising means concentrated leadership; falling means uniform competence. France has the flattest profile in the top forty and falls nine places — excellent everywhere, unbeatable nowhere. Singapore rises on the opposite shape.


8. Reliability, convergence and bias

Convergence. Mean cross-coder SD 1.035 in round 1 → 0.360 after adjudication; exact agreement 6.0 per cent → 37.1 per cent, measured across four adjudicators rather than three. Read the round-1 figure. The coders converge in round 3 because they can see each other, not necessarily because they are more right.

Resolution. Both ranking methods agree at ρ ≈ 0.95 overall, and each adjudicator run alone reproduces 29–30 of the panel’s top thirty. The set is robust; the order inside it is not.

Bias — the panel’s main purpose, and a reassuring null.

Coder Jurisdiction Severity Home country Home region
DeepSeek Flash China −0.081 −0.127 +0.003
Mistral Large France −0.014 +0.334 +0.084
Claude Opus 5 United States −0.034 −0.629 −0.090
DeepSeek V4 Pro China −0.239 +0.271 −0.020
GPT-5.6 Luna United States +0.000 −0.396 −0.003

No coder meaningfully flatters its own jurisdiction. The two positive home lifts are Mistral on France and V4 Pro on China, both under +0.35; both US-jurisdiction models are markedly harsher on the United States than the panel. Regional severity sits inside ±0.10 everywhere.


9. Known problems

1. Governance & Integrity conflates regime type with capability. Fix this first.

Twenty-eight of the 71 Soft Power capabilities sit in this domain, and China’s ten lowest scores in the entire register are all inside it: independent media 2, civic space 3, fundamental rights 3, human rights institutions 3, legislature 6. Drop the domain and China moves from seventh to third.

Every coder scores it the same way, including both Chinese-vendor models (11.1, 11.2, 12.1, 10.8), so this is the framework, not model bias. A one-party state has no independent legislature by construction; scoring it 6 out of 20 measures its constitution, not its capacity to legislate.

Split it into administrative capacity (civil service, digital government, statistical systems, legal identity, regulatory delivery, crisis machinery) and accountability architecture (courts, media, oversight, rights). The two tell entirely different stories: administrative capacity USA 16.52 v China 14.15, a gap of 2.4; accountability architecture USA 15.50 v China 7.00, a gap of 8.5. Equal group weighting makes this worse, not better, because the two halves now split the domain 50/50 regardless of capability counts.

2. No source documents. Not one cell has evidence attached. Evidence vintage is asserted to the coders as 19 September 2026 and is not independently verified.

3. The seam gate is unproven. Pairwise correlation across capabilities within a domain runs high. Some is real; some is halo. Independent coders from different families converging on the same structure weakens the halo explanation without eliminating it. Treat within-domain rank order as firmer than across-domain profile shape.

4. Thin records. Confidence, trajectory and the written assessment record exist only for a sample capability, not the whole register (src/run_records.py produces them).

5. Dropped cells. 480 of ~397,000 judgements (0.12 per cent) were skipped mid-batch by a coder. They are recorded as absent and excluded from the panel median, never zero-filled.


10. Repository layout

ginc-index/
├── README.md                    this file
├── CLAUDE.md                    project instructions for Claude Code
├── .env.example                 API key template
├── src/
│   ├── ginc_paths.py            where the project keeps things; every script imports it
│   ├── all_caps.py              the 288-capability register + per-domain seam text
│   ├── build_caps.py            regenerates all_caps.py from frameworks/*.json
│   ├── nations.csv              the canonical 197 nations (iso3, name, region, subregion)
│   ├── run_all.py               the coder runner: rounds 1 and 3, all five providers
│   ├── reap_locks.py            releases claim locks whose worker died
│   ├── verify_round.py          flags files with failed batches or suspect zero runs
│   ├── build_data.py            adjudication, aggregation, bias, cost → data_index.json
│   ├── pareto.py                Pareto tiers at overall, dimension and domain level
│   ├── build_page.py            builds the dashboard HTML
│   ├── shell.css.html           the dashboard's design system (palette, type, components)
│   ├── oai_bench.py             OpenAI model bench-off used to pick the fifth coder
│   └── run_records.py           structured assessment records (JSONL) for a sample capability
├── frameworks/                  the v2 framework trees the register is built from
├── judgements/                  every coder's scores, one CSV per coder per round per capability
│   └── usage/                   token usage per API call, the source of every cost figure
└── outputs/                     built data, the dashboard, and the published CSVs

File naming in judgements/

Pattern Meaning
r1_<coder>_<slug>.csv round 1, current naming
r3_<coder>_<slug>.csv round 3, current naming
ct1_/ct3_<coder>_<slug>.csv Critical Technology run, earlier naming
<coder>_<slug>.csv the five Energy Technologies pilot capabilities, earliest naming

build_data.py resolves all three schemes. One slug was shortened in the earlier run and is aliased: advanced-therapeutics-vaccines-medical-countermeasures → advanced-therapeutics-vaccines-countermeasures.


11. Running it

Every script imports src/ginc_paths.py, which resolves judgements/, outputs/ and frameworks/ relative to the repository root. Run the scripts from src/; nothing depends on the current working directory. GINC_DATA and GINC_OUT override the two data roots if you keep the judgements elsewhere.

cp .env.example .env          # add your keys, then: chmod 600 .env
set -a; . ./.env; set +a
cd src

# 1. regenerate the register from the framework JSON (only if the frameworks change)
python3 build_caps.py ../frameworks

# 2. code a round. Workers claim capabilities atomically, so run as many as the
#    provider's rate limit allows. 'all' = the whole register; omit for TODO only.
python3 run_all.py 1 coder6 0 1 all      # round, coder, shard, nshards, pool
python3 run_all.py 3 coder6 0 1 all      # round 3 waits until REQUIRE have coded round 1

# 3. verify before trusting anything
python3 verify_round.py 1 coder6          # add --fix to delete and requeue bad files
python3 reap_locks.py                     # release claims whose worker died

# 4. rebuild
python3 build_data.py && python3 pareto.py && python3 build_page.py

Concurrency observed to be safe: DeepSeek 60 workers, OpenAI 24, Anthropic 8, Mistral 6–8 (above that it returns 429 and batches fail silently — always run verify_round.py after Mistral).

Operational traps, all hit at least once:


12. Cost

Model Cost Share
claude-opus-5 $74.26 71%
deepseek-v4-pro $22.10 21%
gpt-5.6-luna $3.91 4%
mistral-large-latest $2.83 3%
deepseek-flash $1.28 1%
Total $104.39

DeepSeek Flash, Mistral Large and GPT-5.6 Luna cost $8.02 between them for the entire register across both rounds. A repeat run of the current five-coder design costs about $104; a three-vendor design using only those three costs about $8 and, on the bias and convergence evidence, loses very little.

Per capability: $0.36 all-in, $0.028 for the three cheap seats. Per rated cell: $0.0018.


13. Next steps, in priority order

  1. Split Governance & Integrity into administrative capacity and accountability architecture. Nothing should be published about China until this is done.
  2. Attach evidence. At minimum, the indicator seed sets already in the register, resolved to values with a source and a date.
  3. Extend the structured assessment record (run_records.py) from a sample to the whole register, so every cell carries a rationale, decisive evidence, the ladder question it failed, and upward and downward triggers.
  4. Human adjudication on the cells where the panel spreads three points or more — 2.7 per cent of the register, about 1,500 cells.
  5. Re-run round 3 symmetrically if a coder is added. GPT-5.6 Luna adjudicated having seen all five round-1 sets; the other adjudicators saw four, because they ran before Luna existed. Cost to make the panel symmetric: about $28, almost all of it DeepSeek V4 Pro.

The data behind it

Everything above is reproducible from analysis/ in this repository: the pipeline, the framework trees, every coder's raw judgements and the built outputs. The published tables are also on this site as pages —capabilities, capability groups,domains, dimensions andnations — and the framework's release archive is atversions.