benchmarks evals
Healthcare Agent Cards Exposed Missing Governance
AGENT-O scored 279 papers against a semantic profile for runtime, tools, provenance, evaluation and governance.
Summary
AGENT-O scored 279 papers against a semantic profile for runtime, tools, provenance, evaluation and governance.
The ontology-based assessment found incomplete reporting in 84.6 percent of papers for runtime architecture, 82.8 percent for governance and safety, and 78.1 percent for provenance and reproducibility. The framework measures reporting completeness, not agent quality or deployment readiness.
Why it matters
AGENT-O scored 279 papers against a semantic profile for runtime, tools, provenance, evaluation and governance.
Limits and context
- The framework measures reporting completeness, not agent quality or deployment readiness.
Key claims
AGENT-O scored 279 papers against a semantic profile for runtime, tools, provenance, evaluation and governance.
Qualification: The framework measures reporting completeness, not agent quality or deployment readiness.
Evidence: source-2026-08-31-020
Sources
- arXiv preprint 2608.28345arXiv · primary research
Corrections
No corrections have been recorded for this story.