Production AI Institute · Public record
Immutable benchmark edition

State of Agent Readiness: August 2026

A frozen PAI Lab edition of visible production-readiness evidence across public AI agent repositories. Use this URL for citations, reporting, and longitudinal comparison.

Repositories24
Average coverage51%
Human oversight42%
Observability75%
Edition finding

Deployment controls were visible; autonomy governance was not yet routine.

In this edition, the strongest visible domain was D5 Deployment control (92% of scanned repositories showed at least one signal). The weakest visible domain was D6 Human oversight (42%). This means maintainers can often show build and release discipline, but many still need explicit approval gates, eval evidence, and operational incident procedures.

Domain coverage

D179%
D258%
D354%
D475%
D592%
D642%
D775%
D858%

What changed into action

  • Publish a docs/production-ai-readiness.md evidence map that links each PSF domain to the relevant repository artifact.
  • Add an evals folder with regression cases, scoring thresholds, and the current model or prompt version under test.
  • Document human approval gates for high-impact actions, including who approves, what is logged, and what fails closed.
  • Add an incident response note, degraded-mode path, and provider fallback procedure.

Findings

  • D5 Deployment control was the strongest visible domain at 92% of scanned repositories.
  • D6 Human oversight was the weakest visible domain at 42% of scanned repositories.
  • Average visible evidence coverage across the sample was 51%.
  • 10 of 24 repositories showed human oversight evidence.

Top visible evidence examples

Evidence pack builder
wlsdks/muse-agentAI observability instrumentation, Human approval gates, Security policy and secret hygiene
83%
IHUI-INF-AI/IHUI-AIAI observability instrumentation, Human approval gates, Security policy and secret hygiene
82%
code-yeongyu/senpiAI observability instrumentation, Human approval gates, Security policy and secret hygiene
77%
7df-lab/devoAI observability instrumentation, Security policy and secret hygiene, Provider fallback or degraded mode
72%
makecindy/cindyAI observability instrumentation, Human approval gates, Security policy and secret hygiene
71%
raydeStar/sir-thaddeusAI observability instrumentation, Security policy and secret hygiene, Provider fallback or degraded mode
65%
vm0-ai/vm0AI observability instrumentation, Security policy and secret hygiene, Provider fallback or degraded mode
62%

Source and method

Repository discovery used GitHub public repository search. The scanner reviewed public repository metadata and public file-tree paths only. It did not clone repositories, inspect private code, or certify the projects listed here.

  • topic:ai-agent archived:false fork:false stars:>=5
  • topic:agentic-ai archived:false fork:false stars:>=5
  • topic:llm-agent archived:false fork:false stars:>=5
  • topic:mcp-server archived:false fork:false stars:>=5
  • "ai agent" in:name,description,readme archived:false fork:false stars:>=5

Citation note

Production AI Institute. "State of Agent Readiness - August 2026." Frozen August 1, 2026. Available at https://www.productionai.institute/agent-readiness/benchmark/2026-08.
  • This edition is an immutable archival snapshot. It does not change when the live benchmark changes.
  • The scan uses public repository metadata and public file-tree paths. It does not clone private code and it is not a certification or endorsement.
  • Later editions are published against the same public methodology and remain fixed after publication.