Model Cards

What if you could explore AI’s frontier through a set of trading cards? Compare models against each other, assess their safety evidence, and discover gaps in their reporting.

Free Systems First Edition trading card packJoin the waitlist ↗
Model index28 cards
01 / 28OpenAI

CAP is our subjective overall rating—not a benchmark score. About CAP

OpenAINo. 028

GPT-6 Astra

OpenAI · GPT-6 · Sep 2026

ARC-AGI-3

99.9

FrontierMath Tier 4 v2

97.6

Terminal-Bench Science 0.1

64.6

Agents’ Last Exam

59.3

AutomationBench

41.4

Advanced math and learning unfamiliar games are defining results of its launch. It also completes scientific workflows, though its business-process tests still leave substantial work unfinished.

Strengths & limitations

Strength

Interactive reasoning

Weakness

End-to-end automation

28 cards

Working edition · September 2026About this set ↗

Model comparison

Model comparison

The evidence landscape

What the published tests measure

541 distinct tests
Share of named tests100%
Counting & sources

Each named test has one primary category. Across all labs, a shared test counts once. A lab’s percentage uses its own inventory as the denominator. Safety includes misuse capabilities, alignment and oversight; these shares describe reporting, not model safety.

Data, coverage & definitions

Named tests and evaluation results use different counting units. A test counts once in the sharing analysis; the evaluation count can include its separately reported settings and subsets.

The inventory combines release tables with safety, capability and oversight assessments from system cards. The source pass now includes the long Anthropic and OpenAI reports and Kimi’s full technical report. All 23 canonical model entries have reconciled text-and-figure result counts. Named-test equivalence and testing conditions are assessed separately.

Dates are model release months, not document revision dates. Points represent publications or target-model counts, not a continuous model lineage. Axes stay fixed when filtering labs. A publication covering multiple models is represented by one word-count point, labeled with both names. Evaluation counts remain model-specific. Source texts include PDF and HTML publications.

Most shared tests counts distinct reporting labs, not matching test conditions. Explicit versions remain separate. Internal tests retain their lab identity. Testing-condition completeness, uncertainty estimates and independent replication still require a systematic review, so no percentages for those fields are shown.

ModelReleaseDocument wordsReported results
o3Apr 20259,914178
GPT-5Aug 202518,968170
GPT-5.1 ThinkingNov 20251,34520
GPT-5.2 ThinkingDec 20256,43372
Gemini 3.1 ProFeb 20261,73430
GPT-5.3 CodexFeb 202610,01232
GPT-5.4Mar 202610,69287
Claude Opus 4.7Apr 202660,962482
GPT-5.5Apr 202613,314113
Claude Opus 4.8May 202661,414566
Mistral Medium 3.5May 2026UnavailableUnavailable
Claude Fable 5Jun 202681,280111
Claude Sonnet 5Jun 202634,007243
Claude Opus 5Jul 202648,174357
GPT-5.6 SolJul 202620,386155
InklingJul 20261,42025
Inkling-SmallJul 20261,49633
Kimi K3Jul 20261,91650
DeepSeek V4-ProAug 2026UnavailableUnavailable
Grok 4.6Aug 20266,22549
Muse GlimmerAug 20262,10730
Muse Spark 1.2Aug 2026UnavailableUnavailable
Claude Fable 5.1Sep 202655,405142
Claude Mythos 5.1Sep 202655,405263
Gemini 3.8 FlashSep 20261,21620
GPT-6 AstraSep 202628,004265
Muse Spark 1.3Sep 2026UnavailableUnavailable

About this set

Welcome to the first edition of the Free Systems Model Cards.

The explorer brings together roughly 476,000 words across 22 system and model cards, alongside supporting release pages and evaluation reports.

You can browse models, compare the benchmarks they share, see where labs are testing different things, trace claims back to their sources, and get a better sense of where the published evidence is genuinely comparable.

The physical deck captures the same snapshot of the frontier in collectible form.

All proceeds go back into the Free Systems community fund to support open, empirical research on AI and politics.

Join the First Edition waitlist ↗

Sources, counting and comparison

Methodology

We use two evidence sets. Official capability sources determine which benchmark results can appear on a card. A narrower corpus of dedicated system and model cards supports comparisons of disclosure length and reporting practice.

Sources

Benchmark results come from first-party system cards, model cards, launch pages and evaluation reports. Each result is stored with its model, benchmark name and version, score, source location and reported test conditions.

Dedicated system and model cards form the canonical disclosure corpus. Launch pages and evaluation reports can support a capability result, but they do not count towards system-card length, evaluation totals or canonical-card overlap.

Benchmark results

A result appears only when the score can be located in an official source, assigned to the correct model and tied to the benchmark label and qualifiers reported beside it. Qualifiers such as tool access, effort level, subset and harness are retained because they can materially change a result.

These are published results, including some attributed to external evaluators. Free Systems has not rerun the evaluations.

Model comparisons

Comparisons follow a fixed order. We first look for the same benchmark with a sufficiently aligned setup. We then look for the same benchmark under different conditions, a related benchmark from the same family, and finally different benchmarks covering the same capability category.

The interface labels those relationships as direct comparison, qualified comparison, same benchmark with a different setup, related benchmark or different evidence. Only the first two support a score comparison. The others show how the labs chose to measure a capability; they do not establish which model is better.

Disclosure length

Disclosure length is the word count of the canonical system or model card. PDF counts use extracted document text; HTML counts use the saved article or model-card body and exclude navigation, scripts and page furniture. Each source is tied to a stored document hash so the count refers to a specific version.

A system card covering multiple models retains its full document count for each model. We do not divide a shared document into estimated model-level word counts.

Editorial fields and limitations

CAP is our subjective overall rating—not a benchmark score. We assign it editorially; it is not a standardized measurement or a calculated average of the bars. The creature artwork is illustrative.

The source material is uneven and changes over time. Labs use different benchmarks, versions, scaffolds and reporting conventions; many results cannot be reproduced from the published information alone. This is a working edition, and corrections will be recorded as the corpus is updated.

Series 1 model cards tuck cover

Series 1
Fall 2026

Join the waitlist

Expected price: $15–$20 + shipping. No payment today.

Expected to ship within 2–3 weeks of orders opening. We’ll email you when the first print run is available.

All proceeds go back into the Free Systems community fund.

Questions? fsystems@stanford.edu