Every Verdit Score is the product of five measurable, reproducible evaluation phases. We publish the weights, the data sources, and the exact threshold definitions so you can audit our reasoning.
Five interconnected phases transform raw GitHub telemetry into enterprise-grade deployment intelligence.
We actively sandbox tools to measure deployment health, init latency, and memory bloat — the Bloat Index — to ensure baseline stability across CPU profiles.
Live TelemetryDeep-dive analysis into context windows, tool-calling accuracy, and agentic reasoning capabilities under adversarial prompt conditions.
In ProgressContinuous OWASP threat intel monitoring, dependency scanning, supply-chain provenance, and compliance readiness for SOC 2 and ISO 27001.
Paywall LiveAggregated ranking of AI infrastructure, comparing tools globally to declare the definitive enterprise leaders — the Gartner Magic Quadrant for AI tooling.
In ProgressEnterprise paywall, tier-gated PDF compliance reports, Clerk authentication, and Stripe billing — the full SaaS monetisation layer.
LiveThis scoring pipeline isn't graded by us alone — a finding only moves the public score once 3 independently-signed reports from distinct network nodes agree on the identical outcome.
How the Canary Network verifies this →Different tool categories get different weight profiles. A vector database is judged differently from an MCP server. Select a category to see its weight distribution.
Stable, secure, well-maintained. Safe to deploy in regulated enterprise environments. Passes OWASP scans. Bloat Index acceptable. Compliance PDF auto-generated.
Acceptable for non-critical workloads. May exhibit memory bloat, slow init latency, or GitHub signal gaps. Review Bloat Index before deploying at scale.
Fails one or more critical checks — abandoned repo, OWASP flag, extreme memory bloat, or recurring deploy failures. Not suitable for production use.
We run every tool inside an isolated Docker container and capture peak RAM at init. This is the number GitHub stars can never tell you.
Every phase above runs inside a real, isolated sandbox — which means every run also produces real execution data: successes, failures, and the exact conditions each occurred under.
This becomes a licensable dataset →CISOs don't buy software. They buy evidence. Here is exactly how our evaluation phases produce artefacts that map to SOC 2, ISO 27001, and NIST AI RMF controls.
/api/v1/audit/:slug — returns all evidence artefacts in a single signed JSON envelope