Calibration ledger

Prediction Ledger

Dated, falsifiable claims belong in public only when the author is ready to own them. Each entry below carries a resolution date and criterion, so the record can be checked, not just believed.

Briernot enough data
resolved0
pending13
ambiguous0

Fewer than 5 resolved predictions, so the score shows the machinery, not evidence of calibration yet.

Open predictions

Seeded 2026-07-03: 13 dated, falsifiable predictions authored by Claude Fable 5 (Anthropic's Mythos-class model) during its limited public window, transcribed and validated verbatim. Scoring begins as they resolve.

The official MCP registry (registry.modelcontextprotocol.io) will list more than 10,000 distinct servers by 2027-01-03.

50% confidence · opened 2026-07-03 · resolves by 2027-01-03 · mcp-ecosystem

Resolves: The official MCP registry's public API or web UI reports a total server count above 10,000 on or before the resolve date; a dated archive capture or API response screenshot suffices.

GitHub's 2026 Octoverse report will state that a majority (over 50%) of surveyed or measured developers' code output involved AI assistance.

55% confidence · opened 2026-07-03 · resolves by 2027-01-03 · portfolio-bets

Resolves: The published Octoverse 2026 report (github.blog or octoverse.github.com) contains a headline or body claim that >50% of code written/pushed by developers in its data or survey involved AI assistance; exact framing may vary but the majority claim must be explicit.

By 2027-01-03, at least two of Anthropic, OpenAI, or Google will publicly sell an individual subscription tier priced at $300/month or more that is marketed primarily around autonomous or long-running agent usage.

45% confidence · opened 2026-07-03 · resolves by 2027-01-03 · ai-tooling

Resolves: Public pricing pages (verifiable via Wayback Machine captures dated on or before 2027-01-03) show two such companies each offering an individual-plan tier at >=$300/month whose marketing copy centers agents or autonomous task execution.

By 2027-01-03, at least two frontier labs will report scores on the same third-party agentic terminal or harness benchmark (such as Terminal-Bench) in their official model announcement posts.

70% confidence · opened 2026-07-03 · resolves by 2027-01-03 · agent-harnesses

Resolves: Two distinct frontier labs' official model release blog posts, published on or before the resolve date, each cite a score on one identical third-party agentic/terminal benchmark; the posts are publicly readable.

By 2027-07-03, the MCP specification will include a MUST-level (mandatory) requirement specifically addressing tool-description injection or tool poisoning.

45% confidence · opened 2026-07-03 · resolves by 2027-07-03 · mcp-ecosystem

Resolves: The MCP specification at modelcontextprotocol.io (or its changelog/commit history) contains normative RFC-2119 MUST language whose stated purpose is mitigating tool-description injection or tool poisoning, merged on or before the resolve date.

By 2027-07-03, at least three major coding harnesses (from: Claude Code, OpenAI Codex, Cursor, GitHub Copilot, Windsurf, Devin) will offer documented parallel multi-agent orchestration as a generally available feature.

85% confidence · opened 2026-07-03 · resolves by 2027-07-03 · agent-harnesses

Resolves: Official public documentation for three of the listed products, live or archived on or before the resolve date, describes spawning multiple concurrent agents/subagents on one task as a GA (non-beta, non-waitlist) feature.

By 2028-07-03, a company with more than 1,000 employees will publish a public incident postmortem that attributes a production outage to a change authored by an autonomous coding agent.

55% confidence · opened 2026-07-03 · resolves by 2028-07-03 · agent-harnesses

Resolves: A public postmortem, incident report, or official engineering blog post from a qualifying company explicitly identifies an AI coding agent as the author of the change that caused a production outage; press coverage quoting the company's own attribution also counts.

By 2028-07-03, Anthropic or OpenAI will be publicly reported, via company statement or credible financial press citing company figures, to exceed a $20B annualized revenue run-rate.

90% confidence · opened 2026-07-03 · resolves by 2028-07-03 · portfolio-bets

Resolves: A company announcement, investor letter, or reporting from an established financial outlet (citing company-provided figures) states annualized revenue or run-rate above $20B for either company on or before the resolve date.

By 2028-07-03, a peer-reviewed paper or major-lab technical report will demonstrate a harness/scaffold change producing a 10+ percentage-point gain on a standard agentic coding benchmark while holding the model fixed.

75% confidence · opened 2026-07-03 · resolves by 2028-07-03 · agent-harnesses

Resolves: A publicly available paper (arXiv or peer-reviewed venue) or official lab report shows, in its own tables, the same model scoring >=10 points higher on SWE-bench Verified, Terminal-Bench, or an equivalent standard benchmark under a modified harness/scaffold.

By 2028-07-03, at least three venture-funded companies whose primary product is MCP or agent-tool security scanning will each have publicly announced funding rounds of $10M or more.

60% confidence · opened 2026-07-03 · resolves by 2028-07-03 · mcp-ecosystem

Resolves: Public funding announcements (company posts, Crunchbase, or press) show three distinct companies, each with a primary product described as MCP/agent-tool security or trust scanning, each with a disclosed round >=$10M announced on or before the resolve date.

By 2028-07-03, an open-weights model that runs on a single consumer GPU with 32GB or less of VRAM will exceed 70% on SWE-bench Verified.

55% confidence · opened 2026-07-03 · resolves by 2028-07-03 · ai-tooling

Resolves: A model with publicly downloadable weights appears at >70% on the SWE-bench Verified leaderboard, and its model card or a reproducible community report documents inference (possibly quantized) within 32GB VRAM on a single consumer card.

By 2028-07-03, a publicly listed model will score 90% or higher on the SWE-bench Verified leaderboard.

88% confidence · opened 2026-07-03 · resolves by 2028-07-03 · ai-tooling

Resolves: The official SWE-bench leaderboard (swebench.com or its successor) shows any model entry at >=90.0% on the Verified split on or before the resolve date; archived leaderboard snapshots count.

By 2028-07-03, a software company with 5 or fewer employees will publicly document $10M or more in ARR for a product built primarily with AI coding agents, with coverage in at least two major tech outlets.

50% confidence · opened 2026-07-03 · resolves by 2028-07-03 · portfolio-bets

Resolves: Public statements from the company (blog, interview, or verified social post) claim >=$10M ARR and <=5 employees and primary AI-agent authorship, corroborated by coverage in two established tech publications on or before the resolve date.

Resolved

No resolved predictions yet.

Calibration curve

Calibration curve waits for at least 10 resolved predictions. Below that, the line would look precise and mean very little.

Interactive calibration tools

A JavaScript-enhanced view of the ledger above: a live position marker on each open prediction, a confidence distribution strip, and a Brier scorecard sandbox. The ledger above is already the full content, with or without any of this.

Rules of the ledger

  • Predictions are append-only in spirit. Corrections happen through resolution and evidence, not quiet edits.
  • Ambiguous is a real outcome, excluded from Brier scoring and counted plainly.
  • The Brier score averages confidence error over correct and incorrect resolutions.
  • Git history is the audit trail, which is less cute than vibes and much more useful.