Agent Hygiene Score

A separate dimension from our security audit — this answers a different question: is this MCP server safe to hand an autonomous AI agent? Five scored dimensions, each 0–2, summed to a total out of 10.

5 Dimensions 0-10 Score 14 Servers Scored 🤖 Agent-First

Why a Separate Score?

Our Security Audit measures whether a server is safe to deploy. The Agent Hygiene Score measures whether it's safe to hand to an autonomous AI agent — a higher bar. An agent doesn't read warning labels. It doesn't ask for confirmation. It follows instructions, and if the server gives it too much power or too little structure, things break in ways the agent can't detect.

🛡️ Security Audit

Question: "Is this server secure?"
Scope: Transport encryption, auth, token lifecycle, input validation, data flow, dependencies.
Audience: DevOps, security teams, procurement.
Score: 0–100

🤖 Agent Hygiene Score

Question: "Is it safe to let an AI agent use this?"
Scope: Schema strictness, least-privilege, boundaries, auditability, maintenance.
Audience: AI engineers, agent builders, Claude/Cursor users.
Score: 0–10

Score Tiers

The total score (sum of all five dimensions) maps to one of four tiers. Each tier comes with a concrete recommendation for agent deployment.

8-10
Excellent
🟢

Safe for autonomous agent use with broad permissions. The server has strict schemas, scoped access, documented boundaries, and auditability. You can trust an agent to use this server without constant oversight.

6-7
Good
🔵

Reasonable with scoped permissions. The server is well-built but may lack auditability or have broad credentials. Give the agent only the tools it needs and monitor initial behavior.

4-5
Fair
🟡

Use with caution. Limit tool access and monitor agent behavior carefully. The server may lack documentation, auditability, or fine-grained scoping. Not ideal for fully autonomous workflows.

0-3
Poor
🔴

Not recommended for autonomous agent use. The server has minimal schema validation, no scoping, no boundaries, and no auditability. Manually review every action the agent takes through this server.

The Five Dimensions

Each dimension is scored 0–2 and summed for the total. Low scores in one dimension aren't necessarily disqualifying — context matters. A local-only server scoring 0 on scoping may still be safe if it runs on the user's machine.

📋 Schema Strictness

Does the server declare typed inputs, enums for constrained choices, and required fields in its tool definitions? Strict schemas prevent malformed agent requests from causing unexpected behavior.

2/2
Typed + Enumerated
All tool inputs use strong types. Enums constrain choices where applicable. Required/optional fields are explicit. JSON Schema or language-native type annotations are complete.
1/2
Partially Typed
Some tools have typed inputs but gaps exist. Enums may be missing. Required fields may not be explicitly declared. Acceptable for simpler servers.
0/2
Untyped / Loose
Tool inputs accept arbitrary strings or objects with no validation. No type constraints, no enums. The agent must guess valid inputs.

🔑 Least-Privilege Scoping

Does the server support per-tool or per-resource authorization scopes, or does one credential unlock everything? Fine-grained scoping lets you give an agent only the access it actually needs.

2/2
Per-Tool Scopes
Each tool or toolset can be individually authorized. OAuth scopes, IAM roles, or tool-level RBAC limit what a connected agent can do. The agent can be granted "read repos" without "write issues".
1/2
Single Credential
One API key or token grants access to all tools. No per-tool scoping. Acceptable for local-only servers where the agent runs on the user's machine.
0/2
No Scoping
Broad credentials with no way to limit access. The same token that reads data can also delete it. Dangerous for autonomous agents.

🚧 Declared Boundaries

Has the author published a security policy, documented rate limits, or explicitly stated what the server will not do? Boundaries tell you what happens when things go wrong.

2/2
Documented Policy
The server has a published security policy, rate limiting, and explicit scope documentation. There is a clear statement of what the server won't do (e.g., "this server never writes to your repo").
1/2
Basic Docs
The README documents what the server does and some configuration options, but does not include an explicit security policy or scope boundaries.
0/2
No Boundaries
No documentation about what the server will or won't do. No rate limits. No guardrails mentioned. The agent has no documented constraints.

🔍 Auditability

Does the server log tool calls, support idempotency keys, or provide observability hooks? Without auditability, you cannot tell what an agent did through the server.

2/2
Full Audit Trail
The server logs tool calls, supports idempotency keys, or integrates with observability platforms (OpenTelemetry, CloudTrail). You can reconstruct what the agent did.
1/2
Partial Observability
Some logging or output that lets you infer tool usage, but no structured audit trail. May rely on the MCP client's own logging.
0/2
No Observability
Tool calls are opaque. No logging, no idempotency, no way to tell what the agent did besides checking external side effects.

🛠️ Maintenance Signal

Is the server actively maintained with recent commits, responsive issue handling, and a growing user base? Abandoned servers may have unpatched vulnerabilities.

2/2
Active & Official
Actively maintained by the project team or an official vendor. Recent commits, responsive issues, and significant GitHub activity. Backed by an organization.
1/2
Community Maintained
Maintained by community contributors. Some recent activity but may have slower response times. Not backed by an official organization.
0/2
Stale / Unknown
No recent commits, unresponsive issues, or unknown maintenance status. Security issues may go unpatched.

Methodology

Scores are assigned from public information only — README files, schema definitions, GitHub metadata, and official documentation. No automated scanning, no LLM inference, no access to proprietary code.

📋 What We Read

GitHub README (primary source), tool schema definitions (JSON Schema, type annotations), GitHub API (stars, forks, commit frequency, issue response time), official documentation sites.

🔄 Maintenance Signal

The Maintenance dimension reuses existing verification data from our catalog — GitHub activity metrics, issue responsiveness, and organization backing. We don't re-score this from scratch.

⚠️ Limitations

Static analysis only. We can't verify runtime behavior, undocumented APIs, or internal security measures. A server scoring 10/10 may still have unknown vulnerabilities. This is a starting point, not a guarantee.

🗓️ Assessment Date

All current scores were assessed on August 12, 2026. Servers are re-evaluated periodically. If a server adds OAuth scoping or audit logging after our assessment, the score will be updated in the next review cycle.

Scored Servers

Initial set of 14 flagship servers scored at launch. This list will grow as we evaluate more servers. Click any server to see the full breakdown on its detail page.

How This Relates to the Security Audit

These are complementary, not competing signals. A server can score well on security (proper transport, auth, input validation) but poorly on agent hygiene (no schema strictness, no audit trail). Both scores appear on each server's detail page.

Aspect Security Audit (0-100) Agent Hygiene (0-10)
Primary question Is this server secure to deploy? Is this server safe for an autonomous agent?
Audience Security teams, DevOps, procurement AI engineers, agent builders, end users
Dims covered Transport, auth, tokens, input, data flow, deps Schema, scoping, boundaries, audit, maintenance
Overlap Input validation ↔ Schema strictness Auth method ↔ Least-privilege scoping
Unique signal Dependency health, token lifecycle, data flow Auditability, declared boundaries, agent-ready schemas
Scoring method Manual audit + self-attestation Public README/schema inspection
Assessed servers 20 servers 14 servers (flagship subset)

Want Your Server Scored?

We're expanding coverage. If you maintain an MCP server and want it evaluated, make sure your README includes: tool schema definitions, authentication documentation, security policy, and observability/logging options. Servers with clear documentation score higher.

Need a Custom MCP System?

Configuration & integration for your stack — from tool selection to production deployment. The directory recommends. The consultancy configures.

Get Started →