analysis/README.md in this repository, the same document the pipeline is run
against.
The home table carries a Tier column. It is a Pareto domination tier over the nine domain scores: one nation dominates another when it is at least as good on all nine and better on at least one, within a quarter of a scale point. Strip the nations nobody dominates and you have tier 1; strip them and repeat for tier 2, and so on.
Section 7 below runs Pareto domination too, but never on that vector. It sorts on a domain's capability-group scores for each of the nine domain events, on the three domain scores inside a dimension, and on the three dimension scores for its overall tier. The home page's tier is a tenth cut, computed by this site, and the two do not agree — which is the point of saying so here rather than letting the word “Pareto” imply they are the same number.
| Sorted on | Criteria | Tier 1 | Tiers | Where it appears |
|---|---|---|---|---|
| The nine domain scores | 9 | 11 | 23 | The home table's Tier column |
| The three dimension scores | 3 | 1 | 60 | §7, the overall tier |
The difference is not a disagreement about the data; it follows from the arithmetic. Adding criteria makes domination harder, because a nation now has to win on all nine rather than on three averages that can hide a weak domain. So the nine-domain frontier is wide, at 11 nations rather than one, and there are 23 tiers beneath it rather than 60. The nine-domain tier is the stricter reading of a nation's shape; the three-dimension tier is the sharper discriminator, and §7 is right that neither is a league table.
The same non-dominated sort is used either way, andscripts/build-ratings.py proves it on every build: it re-runs the sort on the three dimension scores and stops the build unless all 197 nations land in exactly the tier analysis/outputs/ginc_pareto_tier_ranking_2026.csv published. The tolerance is §7's own 0.25 scale points, not a figure this site chose.
🇫🇮Finland🇫🇷France🇩🇪Germany🇯🇵Japan🇳🇱Netherlands🇸🇬Singapore🇰🇷South Korea🇸🇪Sweden🇨🇭Switzerland🇬🇧United Kingdom🇺🇸United States
Alphabetically. Nothing separates them: a tier is a set of nations none of which dominates another, so it carries no internal order. Where the home table sorts by Tier it orders the index inside a tier only to give the rows somewhere to sit.
A complete, reproducible rating of 197 nations against 288 capabilities across the three GINC dimensions, produced by a panel of five independent language-model coders from four vendors and three jurisdictions, in two rounds, for a measured $104.39.
Live dashboard: https://claude.ai/artifact/1Zp4dpyYZjiUHiTgrokGWY
These are candidate ratings, not adopted ratings. They are simulated expert codings produced under the GINC protocol by language models, with no human adjudication and no source documents attached to any cell. Round-1 spread, not round-3 agreement, is the honest reliability estimate.
| Capabilities | 288 (Hard Power 148, Soft Power 71, Economic Power 69) |
| Nations | 197 (the GINC master list, ISO3 as the join key) |
| Rated cells | 56,736 |
| Coder judgements | ~397,000 |
| Coders | 5, across 4 vendors and 3 jurisdictions |
| Rounds | 2 (independent coding, then panel adjudication) |
| Measured cost | $104.39 |
| Dimension | Domain | Groups | Capabilities |
|---|---|---|---|
| Hard Power | Critical Technology | 10 | 47 |
| Hard Power | Strategic Infrastructure | 10 | 62 |
| Hard Power | National Security | 10 | 39 |
| Soft Power | Human Capital | 9 | 25 |
| Soft Power | Information & Influence | 6 | 18 |
| Soft Power | Governance & Integrity | 8 | 28 |
| Economic Power | Financial Strength | 6 | 24 |
| Economic Power | Production & Innovation | 7 | 23 |
| Economic Power | Trade & Investment | 7 | 22 |
Built from the v2 framework JSON in frameworks/. The Factbase v1 tree was discarded: several of
its entries were indicators rather than capabilities (R&D Intensity, GDP Scale & Growth) and
several were normatively contested in a way that collides with the protocol’s sixth principle.
Seven bands of three on a 0–20 range.
| Band | Range | Test |
|---|---|---|
| Frontier | 18–20 | Does the world copy this? |
| Advanced | 15–17 | Among the world’s strongest, holds under stress |
| Established | 12–14 | Works at national scale, institutionalised, reliable |
| Transitional | 9–11 | Works unevenly; strong in the capital, thin beyond |
| Foundational | 6–8 | Basic machinery stands but narrow and fragile |
| Nascent | 3–5 | Pilots and first units only |
| Inception | 0–2 | Absent 0, Declared 1, Initiated 2 |
Coders climb from the bottom and stop at the first no; they judge the band first and place the integer within it second. Criterion-referenced throughout: no regional curve, no home-country lift, no halo from national reputation, no credit across capability seams, no size credit, no regime judgement, no recency inflation.
Every capability carries an explicit seam instruction naming what belongs to the neighbouring domain. Critical Technology rates the capacity to originate, design and manufacture; Strategic Infrastructure rates the deployed estate and the capacity to run it; National Security rates organised force. A nation that buys a whole system from abroad and runs it well scores high on infrastructure and low on technology.
| Coder | Model | Vendor | Jurisdiction | Rounds | Cost |
|---|---|---|---|---|---|
| coder2 | deepseek-chat (Flash) | DeepSeek | China | 1, 3 | $1.28 |
| coder3 | mistral-large-latest | Mistral AI | France | 1, 3 | $2.83 |
| coder4 | claude-opus-5 | Anthropic | United States | 1 | $74.26 |
| coder5 | deepseek-v4-pro | DeepSeek | China | 1 | $22.10 |
| coder6 | gpt-5.6-luna | OpenAI | United States | 1, 3 | $3.91 |
Four adjudicate — DeepSeek Flash, Mistral Large, DeepSeek V4 Pro and GPT-5.6 Luna. Claude Opus 5 has no round-3 pass over the full register and codes round 1 only; its judgement is shown to every adjudicator and its bias stays measurable, but it does not decide. Opus carried 82–90 per cent of the bill on the Critical Technology run and, on the bias and convergence evidence, contributed no more signal than the cheaper models.
Panel parity carries no consequence here because the adjudicated score is a rounded mean rather than a median (see §5). Under a median it would: with four adjudicators the two middle scores differ on 20 per cent of cells, and any tie rule shifts the whole index by about 0.105 points, roughly one coder standard deviation.
Four OpenAI models were tested on two capabilities × 197 nations before one was adopted.
| Model | $ / 288 caps, both rounds | MAE vs panel | Bias | Floor use |
|---|---|---|---|---|
| gpt-5.6-luna | $3.90 | 0.71 / 0.84 | −0.05 / −0.46 | panel-normal |
| gpt-5.6-terra | $29.33 | 0.82 / 0.85 | +0.23 / +0.08 | slightly high |
| gpt-5.6-sol | $67.83 | 1.27 / 1.07 | +1.17 / +0.90 | poor |
| gpt-6-astra | $105.20 | 1.59 / 1.61 | +1.52 / +1.49 | fails |
Agreement and protocol discipline get worse as price rises. On judicial independence the panel
places 10 nations in Inception and 39 in Nascent; gpt-6-astra places zero and eight. It refuses
to use the bottom third of the scale while being marginally harsher at the top — pure scale
compression, which is the one thing an index built for discrimination cannot tolerate. Raw test
data in outputs/openai_coder_test_2026.csv.
Round 1 — independent. Each coder scores all 197 nations on a capability without seeing any other coder. Batched 25–50 nations per call.
Round 3 — panel adjudication. Each adjudicator is shown the full panel’s round-1 table for that capability, with region and sub-region, and asked for its own adjudicated score. The prompt instructs it to adjudicate on evidence rather than average mechanically, to check intra-region ordering and cross-region comparability, and explicitly that “convergence is not the goal; a defensible outlier is worth more than a false consensus.”
Round 2 was dropped. On the Energy Technologies pilot, self-refinement moved the panel almost not at all: mean cross-coder SD went 0.98 → 0.98, with essentially all movement arriving in round 3 when coders could see each other.
The adjudicated score for a nation on one capability is the mean of the four adjudicators’ round-3 integers, rounded to the nearest integer — one integer out of 20, exactly as a human panel would record it.
The rule was tested both ways and depends on panel size:
| Adjudicators | Rule | Cells carrying a score no adjudicator gave |
|---|---|---|
| 3 | median | 0 per cent |
| 3 | rounded mean | 4.08 per cent |
| 4 | rounded mean | 1.95 per cent |
| 4 | median | needs a tie rule; the two middle scores differ on 20 per cent of cells |
With three seats the median wins: a rounded mean returned a value nobody had chosen on one cell in twenty-five, including cases like a panel of 4, 5 and 14 averaging to 8 — a band two of the three rejected. With four seats the mean wins: it uses every seat rather than discarding the extremes, it needs no parity tie-break, and the “nobody chose this” rate more than halves.
The cost of the mean is a small loss of outlier robustness on the 2.7 per cent of cells where the panel spreads three points or more. That is why the panel standard deviation travels with every published figure.
capability → capability group → domain → dimension → index
An unweighted mean at each step. Hard, Soft and Economic Power each count for a third; within each dimension every domain counts once; within each domain every group counts once.
A flat mean over all 288 capabilities would instead hand Hard Power 51.4 per cent of the index, Soft Power 24.7 and Economic Power 24.0, purely because Hard Power happens to be cut into more capabilities. Nobody chose that. The fix barely moves the leading group and corrects the tail: North Korea falls 17 places once its hard-power-heavy profile stops being over-counted.
Rebuilding the whole index separately from each adjudicator’s own scores gives a median standard deviation of 0.124. Against that:
| Granularity | Nations within 1 SD of a boundary |
|---|---|
| 7 protocol bands | 20 of 197 (90 per cent unambiguous) |
| 21 steps with +/− modifiers | 45 of 197 |
| Ranks | at ranks 31–60 the median gap between adjacent nations is 0.073 against a coder SD of ~0.19 |
Publish the band. Order alphabetically within it. Never sort by score inside a band — the moment you do, readers quote the order you just disclaimed. Carry ±SD and a boundary flag on every number.
| Band | n | Nations |
|---|---|---|
| Frontier | 1 | United States (18.03) |
| Advanced | 10 | Canada, China, France, Germany, Japan, Netherlands, Singapore, South Korea, Sweden, United Kingdom |
| Established | 17 | |
| Transitional | 28 | |
| Foundational | 40 | |
| Nascent | 74 | |
| Inception | 27 |
The United States sits at 18.03, barely over the Frontier boundary and flagged as within one standard deviation of it. Under the median rule it scored 17.90 and the Frontier band was empty. Treat “the United States is the only Frontier nation” as a statement the method cannot resolve.
An alternative that needs no weighting at all.
Within each of the nine domains a nation is described by its vector of capability-group scores. Nation A dominates B when A is at least as good on every group and better on at least one. Peeling off the non-dominated set repeatedly gives tiers: tier 1 is the frontier, tier 2 what is left once the frontier is removed, and so on. Nations are ranked lexicographically on how many domains they hold at tier 1, then tier 2, then tier 3, and down through every tier — which separates all 197 with no ties. A tolerance of 0.25 scale points applies, without which rounding noise puts dozens of nations in tier 1 everywhere.
The same comparison is run inside each dimension (three domains each) and inside each domain (tier, then nations dominated).
Use it as a diagnostic, not a league table. It cannot see level: North Korea rises a long way on a single mid-tier domain against an overall index near 4. It is fragile to noise where the mean absorbs it. And it depends on a tolerance parameter that moves the table.
What it is genuinely good for: identifying who holds the frontier (17 of 197 nations hold any of the 41 tier-1 seats across nine domains) and reading a nation’s shape rather than its size. Among the top thirty, the correlation between how far a nation’s strongest domain sits above its own mean and how far it rises on the tier rule is +0.58. Rising means concentrated leadership; falling means uniform competence. France has the flattest profile in the top forty and falls nine places — excellent everywhere, unbeatable nowhere. Singapore rises on the opposite shape.
Convergence. Mean cross-coder SD 1.035 in round 1 → 0.360 after adjudication; exact agreement 6.0 per cent → 37.1 per cent, measured across four adjudicators rather than three. Read the round-1 figure. The coders converge in round 3 because they can see each other, not necessarily because they are more right.
Resolution. Both ranking methods agree at ρ ≈ 0.95 overall, and each adjudicator run alone reproduces 29–30 of the panel’s top thirty. The set is robust; the order inside it is not.
Bias — the panel’s main purpose, and a reassuring null.
| Coder | Jurisdiction | Severity | Home country | Home region |
|---|---|---|---|---|
| DeepSeek Flash | China | −0.081 | −0.127 | +0.003 |
| Mistral Large | France | −0.014 | +0.334 | +0.084 |
| Claude Opus 5 | United States | −0.034 | −0.629 | −0.090 |
| DeepSeek V4 Pro | China | −0.239 | +0.271 | −0.020 |
| GPT-5.6 Luna | United States | +0.000 | −0.396 | −0.003 |
No coder meaningfully flatters its own jurisdiction. The two positive home lifts are Mistral on France and V4 Pro on China, both under +0.35; both US-jurisdiction models are markedly harsher on the United States than the panel. Regional severity sits inside ±0.10 everywhere.
1. Governance & Integrity conflates regime type with capability. Fix this first.
Twenty-eight of the 71 Soft Power capabilities sit in this domain, and China’s ten lowest scores in the entire register are all inside it: independent media 2, civic space 3, fundamental rights 3, human rights institutions 3, legislature 6. Drop the domain and China moves from seventh to third.
Every coder scores it the same way, including both Chinese-vendor models (11.1, 11.2, 12.1, 10.8), so this is the framework, not model bias. A one-party state has no independent legislature by construction; scoring it 6 out of 20 measures its constitution, not its capacity to legislate.
Split it into administrative capacity (civil service, digital government, statistical systems, legal identity, regulatory delivery, crisis machinery) and accountability architecture (courts, media, oversight, rights). The two tell entirely different stories: administrative capacity USA 16.52 v China 14.15, a gap of 2.4; accountability architecture USA 15.50 v China 7.00, a gap of 8.5. Equal group weighting makes this worse, not better, because the two halves now split the domain 50/50 regardless of capability counts.
2. No source documents. Not one cell has evidence attached. Evidence vintage is asserted to the coders as 19 September 2026 and is not independently verified.
3. The seam gate is unproven. Pairwise correlation across capabilities within a domain runs high. Some is real; some is halo. Independent coders from different families converging on the same structure weakens the halo explanation without eliminating it. Treat within-domain rank order as firmer than across-domain profile shape.
4. Thin records. Confidence, trajectory and the written assessment record exist only for a
sample capability, not the whole register (src/run_records.py produces them).
5. Dropped cells. 480 of ~397,000 judgements (0.12 per cent) were skipped mid-batch by a coder. They are recorded as absent and excluded from the panel median, never zero-filled.
ginc-index/
├── README.md this file
├── CLAUDE.md project instructions for Claude Code
├── .env.example API key template
├── src/
│ ├── ginc_paths.py where the project keeps things; every script imports it
│ ├── all_caps.py the 288-capability register + per-domain seam text
│ ├── build_caps.py regenerates all_caps.py from frameworks/*.json
│ ├── nations.csv the canonical 197 nations (iso3, name, region, subregion)
│ ├── run_all.py the coder runner: rounds 1 and 3, all five providers
│ ├── reap_locks.py releases claim locks whose worker died
│ ├── verify_round.py flags files with failed batches or suspect zero runs
│ ├── build_data.py adjudication, aggregation, bias, cost → data_index.json
│ ├── pareto.py Pareto tiers at overall, dimension and domain level
│ ├── build_page.py builds the dashboard HTML
│ ├── shell.css.html the dashboard's design system (palette, type, components)
│ ├── oai_bench.py OpenAI model bench-off used to pick the fifth coder
│ └── run_records.py structured assessment records (JSONL) for a sample capability
├── frameworks/ the v2 framework trees the register is built from
├── judgements/ every coder's scores, one CSV per coder per round per capability
│ └── usage/ token usage per API call, the source of every cost figure
└── outputs/ built data, the dashboard, and the published CSVs
judgements/| Pattern | Meaning |
|---|---|
r1_<coder>_<slug>.csv |
round 1, current naming |
r3_<coder>_<slug>.csv |
round 3, current naming |
ct1_/ct3_<coder>_<slug>.csv |
Critical Technology run, earlier naming |
<coder>_<slug>.csv |
the five Energy Technologies pilot capabilities, earliest naming |
build_data.py resolves all three schemes. One slug was shortened in the earlier run and is
aliased: advanced-therapeutics-vaccines-medical-countermeasures →
advanced-therapeutics-vaccines-countermeasures.
Every script imports src/ginc_paths.py, which resolves judgements/, outputs/ and
frameworks/ relative to the repository root. Run the scripts from src/; nothing depends on the
current working directory. GINC_DATA and GINC_OUT override the two data roots if you keep the
judgements elsewhere.
cp .env.example .env # add your keys, then: chmod 600 .env
set -a; . ./.env; set +a
cd src
# 1. regenerate the register from the framework JSON (only if the frameworks change)
python3 build_caps.py ../frameworks
# 2. code a round. Workers claim capabilities atomically, so run as many as the
# provider's rate limit allows. 'all' = the whole register; omit for TODO only.
python3 run_all.py 1 coder6 0 1 all # round, coder, shard, nshards, pool
python3 run_all.py 3 coder6 0 1 all # round 3 waits until REQUIRE have coded round 1
# 3. verify before trusting anything
python3 verify_round.py 1 coder6 # add --fix to delete and requeue bad files
python3 reap_locks.py # release claims whose worker died
# 4. rebuild
python3 build_data.py && python3 pareto.py && python3 build_page.py
Concurrency observed to be safe: DeepSeek 60 workers, OpenAI 24, Anthropic 8, Mistral 6–8
(above that it returns 429 and batches fail silently — always run verify_round.py after Mistral).
Operational traps, all hit at least once:
pkill -f "<pattern>" kills your own shell when the pattern appears in your command line. Use
pgrep -af to list, then kill explicit PIDs, or put launchers in a script and run them by path.setsid nohup python3 -u ... </dev/null & disown, or they die
with the parent shell.verify_round.py exists because of this.max_tokens=8000
and returns nothing. It needs batch 25 and 32,000 tokens.temperature and rename max_tokens to max_completion_tokens.| Model | Cost | Share |
|---|---|---|
| claude-opus-5 | $74.26 | 71% |
| deepseek-v4-pro | $22.10 | 21% |
| gpt-5.6-luna | $3.91 | 4% |
| mistral-large-latest | $2.83 | 3% |
| deepseek-flash | $1.28 | 1% |
| Total | $104.39 |
DeepSeek Flash, Mistral Large and GPT-5.6 Luna cost $8.02 between them for the entire register across both rounds. A repeat run of the current five-coder design costs about $104; a three-vendor design using only those three costs about $8 and, on the bias and convergence evidence, loses very little.
Per capability: $0.36 all-in, $0.028 for the three cheap seats. Per rated cell: $0.0018.
run_records.py) from a sample to the whole
register, so every cell carries a rationale, decisive evidence, the ladder question it failed,
and upward and downward triggers.Everything above is reproducible from analysis/ in this repository: the pipeline, the framework trees, every coder's raw judgements and the built outputs. The published tables are also on this site as pages —capabilities, capability groups,domains, dimensions andnations — and the framework's release archive is atversions.