Why a Separate Score?
Our Security Audit measures whether a server is safe to deploy. The Agent Hygiene Score measures whether it's safe to hand to an autonomous AI agent — a higher bar. An agent doesn't read warning labels. It doesn't ask for confirmation. It follows instructions, and if the server gives it too much power or too little structure, things break in ways the agent can't detect.
🛡️ Security Audit
Question: "Is this server secure?"
Scope: Transport encryption, auth, token lifecycle, input validation, data flow, dependencies.
Audience: DevOps, security teams, procurement.
Score: 0–100
🤖 Agent Hygiene Score
Question: "Is it safe to let an AI agent use this?"
Scope: Schema strictness, least-privilege, boundaries, auditability, maintenance.
Audience: AI engineers, agent builders, Claude/Cursor users.
Score: 0–10
Score Tiers
The total score (sum of all five dimensions) maps to one of four tiers. Each tier comes with a concrete recommendation for agent deployment.
Safe for autonomous agent use with broad permissions. The server has strict schemas, scoped access, documented boundaries, and auditability. You can trust an agent to use this server without constant oversight.
Reasonable with scoped permissions. The server is well-built but may lack auditability or have broad credentials. Give the agent only the tools it needs and monitor initial behavior.
Use with caution. Limit tool access and monitor agent behavior carefully. The server may lack documentation, auditability, or fine-grained scoping. Not ideal for fully autonomous workflows.
Not recommended for autonomous agent use. The server has minimal schema validation, no scoping, no boundaries, and no auditability. Manually review every action the agent takes through this server.
The Five Dimensions
Each dimension is scored 0–2 and summed for the total. Low scores in one dimension aren't necessarily disqualifying — context matters. A local-only server scoring 0 on scoping may still be safe if it runs on the user's machine.
Schema Strictness
Does the server declare typed inputs, enums for constrained choices, and required fields in its tool definitions? Strict schemas prevent malformed agent requests from causing unexpected behavior.
Least-Privilege Scoping
Does the server support per-tool or per-resource authorization scopes, or does one credential unlock everything? Fine-grained scoping lets you give an agent only the access it actually needs.
Declared Boundaries
Has the author published a security policy, documented rate limits, or explicitly stated what the server will not do? Boundaries tell you what happens when things go wrong.
Auditability
Does the server log tool calls, support idempotency keys, or provide observability hooks? Without auditability, you cannot tell what an agent did through the server.
Maintenance Signal
Is the server actively maintained with recent commits, responsive issue handling, and a growing user base? Abandoned servers may have unpatched vulnerabilities.
Methodology
Scores are assigned from public information only — README files, schema definitions, GitHub metadata, and official documentation. No automated scanning, no LLM inference, no access to proprietary code.
📋 What We Read
GitHub README (primary source), tool schema definitions (JSON Schema, type annotations), GitHub API (stars, forks, commit frequency, issue response time), official documentation sites.
🔄 Maintenance Signal
The Maintenance dimension reuses existing verification data from our catalog — GitHub activity metrics, issue responsiveness, and organization backing. We don't re-score this from scratch.
⚠️ Limitations
Static analysis only. We can't verify runtime behavior, undocumented APIs, or internal security measures. A server scoring 10/10 may still have unknown vulnerabilities. This is a starting point, not a guarantee.
🗓️ Assessment Date
All current scores were assessed on August 12, 2026. Servers are re-evaluated periodically. If a server adds OAuth scoping or audit logging after our assessment, the score will be updated in the next review cycle.
Scored Servers
Initial set of 14 flagship servers scored at launch. This list will grow as we evaluate more servers. Click any server to see the full breakdown on its detail page.
How This Relates to the Security Audit
These are complementary, not competing signals. A server can score well on security (proper transport, auth, input validation) but poorly on agent hygiene (no schema strictness, no audit trail). Both scores appear on each server's detail page.
| Aspect | Security Audit (0-100) | Agent Hygiene (0-10) |
|---|---|---|
| Primary question | Is this server secure to deploy? | Is this server safe for an autonomous agent? |
| Audience | Security teams, DevOps, procurement | AI engineers, agent builders, end users |
| Dims covered | Transport, auth, tokens, input, data flow, deps | Schema, scoping, boundaries, audit, maintenance |
| Overlap | Input validation ↔ Schema strictness | Auth method ↔ Least-privilege scoping |
| Unique signal | Dependency health, token lifecycle, data flow | Auditability, declared boundaries, agent-ready schemas |
| Scoring method | Manual audit + self-attestation | Public README/schema inspection |
| Assessed servers | 20 servers | 14 servers (flagship subset) |
Want Your Server Scored?
We're expanding coverage. If you maintain an MCP server and want it evaluated, make sure your README includes: tool schema definitions, authentication documentation, security policy, and observability/logging options. Servers with clear documentation score higher.