Skip to main content

Browse the directory

Showing 30 of 38 resources for "evaluation"
Saved
Active

37 review · 1 limited in this set — compare to see which signals differ.

Trust snapshot

38 results in this view

Claimed
0%(0/38)

2 trust signals differ in this sample: Source provenance, Submitter

Signals differ on Source provenance, Submitter — add entries to compare before you install.

Rollout signal scan

3 rollout risk signals in current results

Biggest gaps: metadata review, package integrity. 2 entries have 2+ required gaps.

12 scanned

Install payload

Install payload is sparse; verify before rollout decisions.

risk

33% (4/12)

Adoption queue

Browse adoption queue · balanced

26/38 visible results are in hold tier and need mitigation before adoption.

ready 0caution 12hold 26

Arize Phoenix MCP Server for Claude

1 blockers: Metadata review

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

mcp/arize-phoenix-mcp-server · trust review · confidence 67%

Context Engineering Agent Skills

1 blockers: Metadata review

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

skills/context-engineering-agent-skills · trust review · confidence 67%

Firefox DevTools MCP

1 blockers: Metadata review

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

mcp/firefox-devtools-mcp · trust review · confidence 67%

Hugging Face Skills

1 blockers: Metadata review

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

skills/huggingface-skills · trust review · confidence 67%

Inspect AI Benchmark Rubric Agent

1 blockers: Metadata review

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

agents/inspect-ai-benchmark-rubric-agent · trust review · confidence 67%

Langfuse Docs MCP Server for Claude

1 blockers: Metadata review

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

mcp/langfuse-docs-mcp-server · trust review · confidence 67%

LangSmith MCP Server for Claude

1 blockers: Metadata review

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

mcp/langsmith-mcp-server · trust review · confidence 67%

Decision confidence

Decision confidence scan · balanced

25/38 results are low-confidence and need review before adoption.

high 0medium 13low 25

Arize Phoenix MCP Server for Claude

Address Metadata review, Package integrity before broader rollout.

medium

54/100

Missing: Metadata reviewMissing: Package integrity

mcp/arize-phoenix-mcp-server · trust review

Context Engineering Agent Skills

Address Metadata review, Package integrity before broader rollout.

medium

54/100

Missing: Metadata reviewMissing: Package integrity

skills/context-engineering-agent-skills · trust review

Firefox DevTools MCP

Address Metadata review, Package integrity before broader rollout.

medium

54/100

Missing: Metadata reviewMissing: Package integrity

mcp/firefox-devtools-mcp · trust review

Hugging Face Skills

Address Metadata review, Package integrity before broader rollout.

medium

54/100

Missing: Metadata reviewMissing: Package integrity

skills/huggingface-skills · trust review

Inspect AI Benchmark Rubric Agent

Address Metadata review, Package integrity before broader rollout.

medium

54/100

Missing: Metadata reviewMissing: Package integrity

agents/inspect-ai-benchmark-rubric-agent · trust review

Langfuse Docs MCP Server for Claude

Address Metadata review, Package integrity before broader rollout.

medium

54/100

Missing: Metadata reviewMissing: Package integrity

mcp/langfuse-docs-mcp-server · trust review

LangSmith MCP Server for Claude

Address Metadata review, Package integrity before broader rollout.

medium

54/100

Missing: Metadata reviewMissing: Package integrity

mcp/langsmith-mcp-server · trust review

Freshness distribution

Mostly fresh with a few aging entries

Median age 55 days; 10 fresh, 2 aging or stale of 12 scanned.

median 55d

Aging

91–180 days

17%

2 entries

Stale

> 180 days

0%

0 entries

Theme distribution

Results center on evaluation

71% of this view shares the top theme. Leading themes: evaluation, observability, tracing.

Focused

54 distinct themes across 24 scanned

Cursor logo

MIT-licensed Agent Skills collection for context engineering, harness engineering, multi-agent architectures, filesystem context, memory systems, tool design, evaluation, hosted agents, and production agent operating loops for Claude Code, Cursor, Codex, and Open Plugins-compatible agent tools.

Level:expertType:capability-packVerified:validated
Safety ✓ Privacy ✓
Hugging Face logo

Official Hugging Face Agent Skills collection for Claude Code, Codex, Cursor, Gemini CLI, and other skills-compatible agents, covering Hub CLI workflows, datasets, model search, Spaces, Gradio, fine-tuning, evaluations, local models, papers, Trackio, ZeroGPU, transformers.js, TRL, and the Hugging Face MCP server.

Level:expertType:capability-packVerified:validated
Safety ✓ Privacy ✓

MCP server for generating academic diagrams, statistical plots, figure packages, and visual evaluations from research context through PaperBanana's multi-agent illustration pipeline.

Hugging Face logo
Hugging Face Datasetsby Hugging Face · submitted by oktofeesh1

Apache-2.0 library for loading, sharing, streaming, inspecting, and preprocessing AI datasets from the Hugging Face Hub or local files.

Hugging Face logo
Hugging Face Evaluateby Hugging Face · submitted by oktofeesh1

Apache-2.0 library for loading, computing, comparing, saving, and sharing evaluation modules for machine learning models and datasets.

Agenta logo
Agentaby Agenta · submitted by oktofeesh1

Open-source LLMOps platform for prompt management, prompt versioning, evaluation, and observability across LLM applications.

DeepEval logo
DeepEvalby Confident AI · submitted by oktofeesh1

Open-source Python framework for unit-testing LLM applications, agents, RAG pipelines, metrics, regression suites, and traces.

Label Studio logo
Label Studioby HumanSignal · submitted by oktofeesh1

Open-source data labeling, annotation, and human-in-the-loop AI evaluation platform for text, images, audio, video, time series, and multimodal datasets.

Ragas logo
Ragasby Vibrant Labs · submitted by oktofeesh1

Open-source evaluation framework for testing RAG systems, prompts, agents, workflows, and other LLM application behavior.

Langfuse logo

Open-source LLM engineering platform for tracing, prompt management, evaluation, metrics, and observability.

Safety · Privacy ✓

Slash command runbook for designing and running prompt evaluations: define tasks, success criteria, golden outputs, regression checks, and privacy-safe reporting using Anthropic test-and-evaluate guidance.

Invocation:/prompt-eval-runbook <feature-or-prompt-name>
Safety ✓ Privacy ✓
Arize Phoenix logo

Open-source observability and evaluation tooling for LLM applications, traces, datasets, and experiments.

Safety · Privacy ·
Braintrust logo

Evaluation, prompt experimentation, logging, and data platform for production AI application development.

Safety · Privacy ✓
LangSmith logo

Observability, evaluation, tracing, and testing platform for LLM applications and agent workflows.

Safety · Privacy ✓
Weave logo

Weights and Biases toolkit for tracking, evaluating, and debugging LLM applications and agent workflows.

Safety · Privacy ·

scikit-learn ML modeling rule that audits for data leakage, enforces Pipeline-based preprocessing, and validates cross-validation rigor (StratifiedKFold, GroupKFold, TimeSeriesSplit) for defensible model evaluation

Safety · Privacy ·
Prompt flow logo
Prompt flowby microsoft · submitted by davion-knight

Open-source suite of development tools from Microsoft for building LLM applications end to end — create executable flows that link LLMs, prompts, Python, and tools, trace and debug them, evaluate quality against datasets in CI/CD, and deploy to a serving platform.

Inspect LLM traces and spans, manage prompts, explore datasets, and review evaluation experiments from Claude — with the official Arize Phoenix MCP server, built into the open-source Phoenix AI observability platform.

LangSmith logo

Connect Claude to LangSmith — retrieve conversation threads and traces, fetch and push prompts, browse evaluation datasets and experiments, and access billing usage — with the official LangSmith Model Context Protocol server from LangChain.

Debug, evaluate, and monitor LLM applications from Claude — read traces and spans, score outputs, save prompts, run evaluation experiments, and query project metrics — with the official Opik MCP server by Comet.

Evidently logo
Evidentlyby Evidently AI · submitted by oktofeesh1

Open-source ML and LLM observability framework for evaluating, testing, and monitoring data quality, drift, model behavior, and AI application outputs.

MLflow logo
MLflowby MLflow Project · submitted by oktofeesh1

Open-source AI engineering platform for tracing, evaluating, prompt-managing, and deploying agents, LLM applications, and ML models.

TruLens logo
TruLensby TruEra / Snowflake · submitted by oktofeesh1

Open-source evaluation and tracing framework for measuring AI agents, RAG systems, LLM apps, retrieval quality, feedback metrics, and trace-level regressions.

Giskard logo

AI testing platform for evaluating, scanning, and monitoring machine learning and LLM application quality.

Safety · Privacy ·
AWS logo

Capability comparison of Claude Code, Amazon Q Developer (formerly CodeWhisperer), and Google Gemini Code Assist: form factor, agentic vs completion, IDE support, cloud ties, and free tiers.

Safety · Privacy ✓
Windsurf logo

Capability comparison of Claude Code, Cursor, and Windsurf (formerly Codeium): form factor, where each runs, agentic vs autocomplete, MCP extensibility, and free tiers, grounded in each tool's official docs.

Safety · Privacy ✓
GitHub Copilot logo

Capability comparison of Claude Code, GitHub Copilot, and ChatGPT (Codex) for Python development. Form factor, where each runs, agentic vs inline, IDE integration, MCP, and free tiers, grounded in each tool's official docs.

Safety · Privacy ✓

Source-backed agent for designing Inspect AI benchmark tasks, datasets, solver plans, scorer rubrics, model matrices, eval logs, and release-quality prompt evaluation decisions.

OpenAI Evalsby OpenAI · submitted by JSONbored

Open-source framework from OpenAI for evaluating LLM and agent behavior with reusable eval definitions, grading logic, datasets, and regression workflows.

A practical guide for comparing AI coding tools with repeatable benchmarks, fixed task sets, controlled environments, transparent scoring, and privacy-safe artifacts.