Skip to main content

Browse the directory

Showing 27 resources for "evaluation"
Saved
Active

Source-backed filter active — add entries to compare trust side by side.

Trust snapshot

27 results in this view

Claimed
0%(0/27)

1 trust signal differs in this sample: Submitter

Signals differ on Submitter — add entries to compare before you install.

Rollout signal scan

3 rollout risk signals in current results

Biggest gaps: metadata review, package integrity. 2 entries have 2+ required gaps.

12 scanned

Install payload

Install payload is sparse; verify before rollout decisions.

risk

25% (3/12)

Adoption queue

Browse adoption queue · balanced

18/27 visible results are in hold tier and need mitigation before adoption.

ready 0caution 9hold 18

Arize Phoenix MCP Server for Claude

1 blockers: Metadata review

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

mcp/arize-phoenix-mcp-server · trust review · confidence 67%

Context Engineering Agent Skills

1 blockers: Metadata review

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

skills/context-engineering-agent-skills · trust review · confidence 67%

Firefox DevTools MCP

1 blockers: Metadata review

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

mcp/firefox-devtools-mcp · trust review · confidence 67%

Hugging Face Skills

1 blockers: Metadata review

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

skills/huggingface-skills · trust review · confidence 67%

Inspect AI Benchmark Rubric Agent

1 blockers: Metadata review

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

agents/inspect-ai-benchmark-rubric-agent · trust review · confidence 67%

LangSmith MCP Server for Claude

1 blockers: Metadata review

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

mcp/langsmith-mcp-server · trust review · confidence 67%

OpenAI Evals

1 blockers: Metadata review

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

tools/openai-evals · trust review · confidence 67%

Opik MCP Server for Claude

1 blockers: Metadata review

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

mcp/opik-mcp-server · trust review · confidence 67%

Decision confidence

Decision confidence scan · balanced

18/27 results are low-confidence and need review before adoption.

high 0medium 9low 18

Arize Phoenix MCP Server for Claude

Address Metadata review, Package integrity before broader rollout.

medium

54/100

Missing: Metadata reviewMissing: Package integrity

mcp/arize-phoenix-mcp-server · trust review

Context Engineering Agent Skills

Address Metadata review, Package integrity before broader rollout.

medium

54/100

Missing: Metadata reviewMissing: Package integrity

skills/context-engineering-agent-skills · trust review

Firefox DevTools MCP

Address Metadata review, Package integrity before broader rollout.

medium

54/100

Missing: Metadata reviewMissing: Package integrity

mcp/firefox-devtools-mcp · trust review

Hugging Face Skills

Address Metadata review, Package integrity before broader rollout.

medium

54/100

Missing: Metadata reviewMissing: Package integrity

skills/huggingface-skills · trust review

Inspect AI Benchmark Rubric Agent

Address Metadata review, Package integrity before broader rollout.

medium

54/100

Missing: Metadata reviewMissing: Package integrity

agents/inspect-ai-benchmark-rubric-agent · trust review

LangSmith MCP Server for Claude

Address Metadata review, Package integrity before broader rollout.

medium

54/100

Missing: Metadata reviewMissing: Package integrity

mcp/langsmith-mcp-server · trust review

OpenAI Evals

Address Metadata review, Package integrity before broader rollout.

medium

54/100

Missing: Metadata reviewMissing: Package integrity

tools/openai-evals · trust review

Opik MCP Server for Claude

Address Metadata review, Package integrity before broader rollout.

medium

54/100

Missing: Metadata reviewMissing: Package integrity

mcp/opik-mcp-server · trust review

Freshness distribution

Mostly fresh with a few aging entries

Median age 55 days; 10 fresh, 2 aging or stale of 12 scanned.

median 55d

Aging

91–180 days

17%

2 entries

Stale

> 180 days

0%

0 entries

Theme distribution

Results center on evaluation

58% of this view shares the top theme. Leading themes: evaluation, observability, tracing.

Focused

62 distinct themes across 24 scanned

Cursor logo

MIT-licensed Agent Skills collection for context engineering, harness engineering, multi-agent architectures, filesystem context, memory systems, tool design, evaluation, hosted agents, and production agent operating loops for Claude Code, Cursor, Codex, and Open Plugins-compatible agent tools.

Level:expertType:capability-packVerified:validated
Safety ✓ Privacy ✓
Hugging Face logo

Official Hugging Face Agent Skills collection for Claude Code, Codex, Cursor, Gemini CLI, and other skills-compatible agents, covering Hub CLI workflows, datasets, model search, Spaces, Gradio, fine-tuning, evaluations, local models, papers, Trackio, ZeroGPU, transformers.js, TRL, and the Hugging Face MCP server.

Level:expertType:capability-packVerified:validated
Safety ✓ Privacy ✓

MCP server for generating academic diagrams, statistical plots, figure packages, and visual evaluations from research context through PaperBanana's multi-agent illustration pipeline.

Hugging Face logo
Hugging Face Datasetsby Hugging Face · submitted by oktofeesh1

Apache-2.0 library for loading, sharing, streaming, inspecting, and preprocessing AI datasets from the Hugging Face Hub or local files.

Hugging Face logo
Hugging Face Evaluateby Hugging Face · submitted by oktofeesh1

Apache-2.0 library for loading, computing, comparing, saving, and sharing evaluation modules for machine learning models and datasets.

Agenta logo
Agentaby Agenta · submitted by oktofeesh1

Open-source LLMOps platform for prompt management, prompt versioning, evaluation, and observability across LLM applications.

DeepEval logo
DeepEvalby Confident AI · submitted by oktofeesh1

Open-source Python framework for unit-testing LLM applications, agents, RAG pipelines, metrics, regression suites, and traces.

Label Studio logo
Label Studioby HumanSignal · submitted by oktofeesh1

Open-source data labeling, annotation, and human-in-the-loop AI evaluation platform for text, images, audio, video, time series, and multimodal datasets.

Ragas logo
Ragasby Vibrant Labs · submitted by oktofeesh1

Open-source evaluation framework for testing RAG systems, prompts, agents, workflows, and other LLM application behavior.

Langfuse logo

Open-source LLM engineering platform for tracing, prompt management, evaluation, metrics, and observability.

Safety · Privacy ✓
Arize Phoenix logo

Open-source observability and evaluation tooling for LLM applications, traces, datasets, and experiments.

Safety · Privacy ·
Prompt flow logo
Prompt flowby microsoft · submitted by davion-knight

Open-source suite of development tools from Microsoft for building LLM applications end to end — create executable flows that link LLMs, prompts, Python, and tools, trace and debug them, evaluate quality against datasets in CI/CD, and deploy to a serving platform.

Inspect LLM traces and spans, manage prompts, explore datasets, and review evaluation experiments from Claude — with the official Arize Phoenix MCP server, built into the open-source Phoenix AI observability platform.

LangSmith logo

Connect Claude to LangSmith — retrieve conversation threads and traces, fetch and push prompts, browse evaluation datasets and experiments, and access billing usage — with the official LangSmith Model Context Protocol server from LangChain.

Debug, evaluate, and monitor LLM applications from Claude — read traces and spans, score outputs, save prompts, run evaluation experiments, and query project metrics — with the official Opik MCP server by Comet.

Evidently logo
Evidentlyby Evidently AI · submitted by oktofeesh1

Open-source ML and LLM observability framework for evaluating, testing, and monitoring data quality, drift, model behavior, and AI application outputs.

MLflow logo
MLflowby MLflow Project · submitted by oktofeesh1

Open-source AI engineering platform for tracing, evaluating, prompt-managing, and deploying agents, LLM applications, and ML models.

TruLens logo
TruLensby TruEra / Snowflake · submitted by oktofeesh1

Open-source evaluation and tracing framework for measuring AI agents, RAG systems, LLM apps, retrieval quality, feedback metrics, and trace-level regressions.

Giskard logo

AI testing platform for evaluating, scanning, and monitoring machine learning and LLM application quality.

Safety · Privacy ·

Source-backed agent for designing Inspect AI benchmark tasks, datasets, solver plans, scorer rubrics, model matrices, eval logs, and release-quality prompt evaluation decisions.

OpenAI Evalsby OpenAI · submitted by JSONbored

Open-source framework from OpenAI for evaluating LLM and agent behavior with reusable eval definitions, grading logic, datasets, and regression workflows.

Mozilla-maintained MCP server for automating Firefox through WebDriver BiDi, with tools for page navigation, snapshots, UID-based input, screenshots, network requests, console messages, dialogs, history, viewport changes, optional JavaScript evaluation, privileged Firefox contexts, preferences, and WebExtension.

Gradio logo
Gradioby Gradio · submitted by oktofeesh1

Apache-2.0 Python framework for building and sharing machine-learning demos, AI web apps, model interfaces, chatbots, API front ends, and interactive evaluation tools.

LlamaIndex logo
LlamaIndexby LlamaIndex · submitted by oktofeesh1

Open-source framework for building agentic LLM applications over private data with ingestion, indexes, retrieval, RAG, tools, workflows, and evaluation.

mini-SWE-agent logo
mini-SWE-agentby SWE-agent · submitted by oktofeesh1

MIT-licensed command-line software-engineering agent for local coding tasks, GitHub issue fixing, trajectory inspection, and SWE-bench style evaluation.

Promptfoo logo

Open-source prompt testing and red-teaming framework for LLM outputs, regressions, evaluations, and security checks.

Safety · Privacy ✓
Mastra logo

TypeScript agent framework for building AI agents, workflows, memory, tool calling, and evaluation-backed applications.

Safety · Privacy ·