Vibe Coding Discover

AI Tools

AgentMeasure

View on GitHub

Open measurement infrastructure for AI agents. Our audit of 124 usage tools found 45+ verified billing bugs — 19 fixes landed upstream. Conformance fixtures for token accounting: PASS / FAIL / UNPROVABLE in CI.

★ 2183 forksPythonMITroy-tong

Open measurement and conformance layer for AI-agent telemetry. Audits token and usage accounting with PASS/FAIL/UNPROVABLE CI checks, reads Codex rollout logs locally to expose retries and duplicate records, and runs preregistered harness experiments.

Use Cases

Verify token/usage accounting labels in CI with PASS/FAIL/UNPROVABLE conformance fixturesDetect retry inflation and duplicate records in Codex session rollout logs locallySeparate attempts (execution facts) from logical operations in agent reliability reportingRun preregistered task-set x harness A/B experiments with budget circuit breakersAudit third-party usage tools for billing bugs and report evidence-backed findingsGenerate local HTML measurement reports without sending raw logs off-machineBenchmark competing agent harnesses across reach, choice, success and consumption funnelsDefine standard billable outcome units (resolutions, completed tasks) for agent providers

Built With

Language
Python
Frameworks
OpenTelemetry · MCP · Pydantic AI · GitHub Actions · pipx

Tags

agent-observability · token-accounting · telemetry · conformance-testing · llm-metrics · usage-tracking · opentelemetry · local-first · agent-evaluation · mcp · codex · claude-code · ci-checks · benchmarking · standards · quality-ladder

AgentMeasure — Vibe Coding Discover