Scoring Standard

No Black Boxes.
Full Transparency.

Every Verdit Score is the product of five measurable, reproducible evaluation phases. We publish the weights, the data sources, and the exact threshold definitions so you can audit our reasoning.

· · 5 Evaluation Phases · Updated 2026-06-28
Proprietary Evaluation Standard

The Industry Standard for AI Tool Evaluation.

Five interconnected phases transform raw GitHub telemetry into enterprise-grade deployment intelligence.

Phase 01
Foundation

We actively sandbox tools to measure deployment health, init latency, and memory bloat — the Bloat Index — to ensure baseline stability across CPU profiles.

Live Telemetry
Phase 02
Expertise

Deep-dive analysis into context windows, tool-calling accuracy, and agentic reasoning capabilities under adversarial prompt conditions.

In Progress
Phase 03
Maturity

Continuous OWASP threat intel monitoring, dependency scanning, supply-chain provenance, and compliance readiness for SOC 2 and ISO 27001.

Paywall Live
Phase 04
Index

Aggregated ranking of AI infrastructure, comparing tools globally to declare the definitive enterprise leaders — the Gartner Magic Quadrant for AI tooling.

In Progress
Phase 05
Monetisation

Enterprise paywall, tier-gated PDF compliance reports, Clerk authentication, and Stripe billing — the full SaaS monetisation layer.

Live

This scoring pipeline isn't graded by us alone — a finding only moves the public score once 3 independently-signed reports from distinct network nodes agree on the identical outcome.

How the Canary Network verifies this →
Signal Architecture

Category-Adaptive Signal Weights.

Different tool categories get different weight profiles. A vector database is judged differently from an MCP server. Select a category to see its weight distribution.

GitHub Stars
20%
Issue Activity
15%
Commit Recency
15%
Deploy Health
20%
Bloat Index (RAM)
15%
OWASP Clean
10%
License Clarity
5%
Score Tiers

What the Score Actually Means.

Production Grade

80+

Stable, secure, well-maintained. Safe to deploy in regulated enterprise environments. Passes OWASP scans. Bloat Index acceptable. Compliance PDF auto-generated.

✓ DEPLOY — RECOMMENDED

Conditional

60–79

Acceptable for non-critical workloads. May exhibit memory bloat, slow init latency, or GitHub signal gaps. Review Bloat Index before deploying at scale.

⚠ DEPLOY — WITH REVIEW

Not Recommended

0–59

Fails one or more critical checks — abandoned repo, OWASP flag, extreme memory bloat, or recurring deploy failures. Not suitable for production use.

✗ DO NOT DEPLOY
Phase 01 Detail

The Bloat Index — Real Numbers.

We run every tool inside an isolated Docker container and capture peak RAM at init. This is the number GitHub stars can never tell you.

Lowest Bloat (Recommended)
12.4 MB
OB1 · Database · 3,645 ★
Highest Measured Bloat
99.4 MB
SwarmClaw · AI/ML · 565 ★
Signal Quality
99.1%
Harness bugs excluded · 15,196 runs
Tools Exceeding 500MB Threshold
0
Across all 112 indexed tools

Every phase above runs inside a real, isolated sandbox — which means every run also produces real execution data: successes, failures, and the exact conditions each occurred under.

This becomes a licensable dataset →
Enterprise Trust Architecture

Methodology Mapped to Compliance Frameworks.

CISOs don't buy software. They buy evidence. Here is exactly how our evaluation phases produce artefacts that map to SOC 2, ISO 27001, and NIST AI RMF controls.

SOC 2
CC6.1Logical & Physical AccessAuth-header enforcement verified via Ghost Swarm probe on all MCP endpoints
CC7.2System MonitoringContinuous sandbox telemetry retained in D1 for 90 days, exportable on demand
CC9.2Risk Assessment12-signal Verdit Score provides ongoing risk quantification per tool
Artefacts Generated
📄
Signed PDF report per tool — includes sandbox log hash, OWASP result, and Bloat Index with timestamp
🔗
D1 immutable ledger entry per run — block number, row hash, and signal quality score
📋
Auditor-ready export via /api/v1/audit/:slug — returns all evidence artefacts in a single signed JSON envelope
ISO 27001
A.8.8Vulnerability ManagementCVE cross-reference against SBOM at every sandbox run via OSV advisory database
A.8.12Data Leakage PreventionOWASP LLM06 payload injection tests probe for outbound data exfiltration paths
A.8.28Secure CodingDependency tree extracted from package manifests, supply-chain provenance verified
Artefacts Generated
🔬
CycloneDX SBOM — signed, versioned, includes all transitive dependencies and known CVEs
📊
Supply-chain risk score — derived from bus factor, maintainer count, and fork depth
NIST AI RMF
GOVERN 1.1Risk PolicyVerdit Score tier provides a deployability policy signal organisations can adopt directly
MEASURE 2.5AI Risk MetricsBloat Index, pass@k rate, and OWASP count provide quantified AI risk metrics per tool
MANAGE 2.4Incident ResponseLive threat feed and score-drop webhooks provide real-time SOC-digestible telemetry
EU AI Act — Q3 2026
🇪🇺
Article 9 risk management and Annex IV technical documentation generation in active development. Enterprise waitlist available.
Get Signed Compliance PDF → Pro plan required · Same-day delivery · Auditor-ready format