Methodology
Grade
Start at 100. Subtract 25 per open blocker, 5 per warning and 1 per info finding; the floor is 0. A ≥ 90, B ≥ 80, C ≥ 70, D ≥ 60, otherwise F. Any open blocker caps the grade at D, and any open security blocker caps it at F. Suppressed findings do not count.
Precision
Every rule is measured on a benchmark of 210 public repositories that contain agent context files, selected with GitHub code search on 2026-10-03 across AGENTS.md, CLAUDE.md, .cursorrules, .cursor/rules, Copilot instructions, GEMINI.md and .mcp.json, each pinned to a commit. Every finding was reviewed by hand and labelled true or false positive.
- A deterministic rule becomes GA with at least 5 corpus findings and under 5% false positives.
- A rule at or over 5% is demoted automatically: it reports at info and stays beta.
- A rule with too few findings to measure stays beta.
The corpus is selected for repositories that already have context files, so it says nothing about the share of all repositories. Findings for individual public repositories are never published.
Declared, not observed
threadctx shows what a correctly behaving tool loads according to its documentation. It does not observe a running agent. The Claude Code adapter is additionally checked against the real tool.
Token estimates
Counts are estimates (about four characters per token) and always labelled as such.