Results are only as reliable as the evidence behind each answer. Teams may score intended policy instead of observed practice. Use a second reviewer and inspect artifacts for consequential decisions.
Limitations
This is a decision aid, not a certification, compliance opinion, security assessment, or promise of agent performance.
Self-assessment bias
Context matters
A safe operating pattern depends on data sensitivity, reversibility, regulatory context, deployment surface, team skill, and failure impact. The same score can represent different residual risks.
Not a benchmark
Bands are transparent operating thresholds, not industry percentiles. This version has not been normed against a representative population or validated as a predictive psychometric instrument.
Recommendation scope
The three recommended actions are a deterministic starting point, not an exhaustive remediation plan. They select one action from each of the three highest-priority dimensions; when fewer than three dimensions contain a score deficit, an already-strong dimension can still appear after every actual gap has been prioritized.
Model capability is out of scope
The audit scores the system around agentic coding. It does not rank models or guarantee correctness, security, accessibility, maintainability, or return on investment.
Use qualified review
Security, privacy, legal, compliance, labor, and safety decisions require qualified domain review. A Controlled result does not remove that responsibility.