Limitations

This is a decision aid, not a certification, compliance opinion, security assessment, or promise of agent performance.

Self-assessment bias

Results are only as reliable as the evidence behind each answer. Teams may score intended policy instead of observed practice. Use a second reviewer and inspect artifacts for consequential decisions.

Context matters

A safe operating pattern depends on data sensitivity, reversibility, regulatory context, deployment surface, team skill, and failure impact. The same score can represent different residual risks.

Not a benchmark

Bands are transparent operating thresholds, not industry percentiles. This version has not been normed against a representative population or validated as a predictive psychometric instrument.

Recommendation scope

The three recommended actions are a deterministic starting point, not an exhaustive remediation plan. They select one action from each of the three highest-priority dimensions; when fewer than three dimensions contain a score deficit, an already-strong dimension can still appear after every actual gap has been prioritized.

Model capability is out of scope

The audit scores the system around agentic coding. It does not rank models or guarantee correctness, security, accessibility, maintainability, or return on investment.

Use qualified review

Security, privacy, legal, compliance, labor, and safety decisions require qualified domain review. A Controlled result does not remove that responsibility.