Docs

SwarmBench is a measurement surface, not an API product. Everything it publishes is a static, hashed artifact you can fetch and verify. Note: this surface previously described a "BFT consensus for MoE routing" library — that positioning was retired after the n_eff 1.21 refutation. This page replaces those docs.

/leaderboard/data.json

The measured leaderboard as JSON: 58 models, per-model practice/held-out accuracy, overfit gap, tokens-per-correct, cross-substrate deltas and flags, plus sha256 of every source artifact. Regenerated by the daily lane.

/methodology

Protocol v1: frozen splits, three-outcome honesty, n_eff accounting, panel aggregation, cross-substrate rules, salt-split reproducibility.

/llms.txt

Plain-text facts for LLM consumers: what SwarmBench measures, the six metrics, the API surface.

/agent.json + /.well-known/agent-card.json

A2A discovery: what SwarmBench measures and how an agent submits a mixture.

Verify (free, no key)

https://councilof.ai/verify-server/ — look up signed measurements about an endpoint.
https://councilof.ai/mcp — MCP server with free verification tools such as verify_card and verify_capsule.
https://councilof.ai/api/a2a — A2A endpoint.