Claude resources tagged “evaluation”
20 curated Claude Code resources tagged evaluation in the HeyClaude directory — mostly tools, mcp servers, and guides.
Highlights from this set
Standout entries tagged evaluation, picked by their own metadata — trust tier, provenance, documentation, and recency.
All evaluation resources
Arize Phoenix
Open-source LLM observability and evaluation tooling.
Arize Phoenix MCP Server for Claude
Official Arize Phoenix MCP: LLM traces, prompts, datasets & experiments from Claude.
Braintrust
Evaluation and logging platform for AI applications.
Claude Code vs Amazon Q Developer vs Gemini Code Assist
Compare Claude Code, Amazon Q Developer (formerly CodeWhisperer), and Google Gemini Code Assist on form factor, agentic features, and ecosystem fit.
Claude Code vs Cursor vs Windsurf (Codeium)
Compare Claude Code, Cursor, and Windsurf (formerly Codeium) on form factor, agentic features, MCP support, and free tiers.
Claude Code vs GitHub Copilot vs ChatGPT for Python Dev
How Claude Code, GitHub Copilot, and ChatGPT (Codex) compare for Python work.
Context Engineering Agent Skills
Agent Skills pack for context windows, multi-agent systems, filesystem memory, tool design, evaluation, harness engineering, and production agents.
DeepEval
Python LLM evaluation tests, metrics, regression checks, and tracing.
Evidently
Evaluate, test, and monitor ML models, LLM apps, data quality, and drift.
Hugging Face Evaluate
Load metrics, comparisons, and measurements for reproducible model and dataset evaluation.
Label Studio
Open-source data labeling and human-in-the-loop AI evaluation.
LangSmith
Observability and evaluation platform for LLM apps.
LangSmith MCP Server for Claude
Official LangSmith MCP: traces, prompts, datasets & experiments from Claude.
MLflow
Trace, evaluate, monitor, and manage agents, LLM apps, prompts, and models.
Opik MCP Server for Claude
Official Opik MCP: LLM traces, evals, prompts & experiments from Claude.
Prompt flow
Open-source Microsoft toolkit to build, trace, evaluate, and deploy LLM app flows.
Ragas
Open-source evaluation framework for RAG and LLM application testing.
TruLens
Evaluate and trace agents, RAG systems, LLM apps, metrics, and regressions.
Signal coverage
How these 20 resources score on the trust and safety signals HeyClaude reviews — counted from this set, not the directory as a whole.
A short, calm digest of reviewed Claude resources. Unsubscribe any time.