Claude resources tagged “evals”
10 curated Claude Code resources tagged evals in the HeyClaude directory — mostly mcp servers, guides, and tools. 1 of them sits in the trusted tier.
Highlights from this set
Standout entries tagged evals, picked by their own metadata — trust tier, provenance, documentation, and recency.
All evals resources
/prompt-eval-runbook - Prompt Eval Runbook Slash Command
Run a structured prompt evaluation runbook with tasks, criteria, and regression checks.
Agent Evals Regression Gate Skill
Create a practical eval harness with golden sets, rubric scoring, and release gates for agentic workflows.
Arize Phoenix MCP Server for Claude
Official Arize Phoenix MCP: LLM traces, prompts, datasets & experiments from Claude.
Evaluate AI Coding Tools with Repeatable Benchmarks
Compare AI coding tools with repeatable benchmark runs.
Laminar
Open-source AI-agent observability with tracing, evals, signals, SQL dashboards, and datasets.
Langfuse Docs MCP Server for Claude
Public Langfuse Docs MCP server for agent-ready documentation search, page retrieval, and implementation guidance over Langfuse observability docs.
Open Source Evals Prompt Testing
Open-source evals bundle for prompt tests, RAG metrics, traces, human review, and regression gates.
OpenAI Agents Trace to Eval Regression Guide
Turn OpenAI agent traces into repeatable regression evals.
Signal coverage
How these 10 resources score on the trust and safety signals HeyClaude reviews — counted from this set, not the directory as a whole.
A short, calm digest of reviewed Claude resources. Unsubscribe any time.