Skip to main content

Browse the directory

Showing 16 resources for "regression"
Saved
Active

2 trusted · 13 review · 1 limited in this set — compare to see which signals differ.

Trust snapshot

16 results in this view

Claimed
0%(0/16)

3 trust signals differ in this sample: Package trust, Source provenance, Submitter

Signals differ on Package trust, Source provenance, Submitter — add entries to compare before you install.

Rollout signal scan

2 rollout risk signals in current results

Biggest gaps: metadata review, package integrity. 0 entries have 2+ required gaps.

12 scanned

Install payload

Install payload is broadly covered in current results.

good

83% (10/12)

Adoption queue

Browse adoption queue · balanced

3/16 visible results are in hold tier and need mitigation before adoption.

ready 0caution 13hold 3

Agent Evals Regression Gate Skill

1 blockers: Metadata review

caution

70/100

Request metadata review from maintainers or internal owners.

skills/agent-evals-regression-gate · trust trusted · confidence 83%

MCP Tool Contract Testing Skill

1 blockers: Metadata review

caution

70/100

Request metadata review from maintainers or internal owners.

skills/mcp-tool-contract-testing · trust trusted · confidence 83%

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

skills/claude-code-troubleshooting-triage-capability-pack · trust review · confidence 67%

Codecov Patch Coverage Planning Agent

1 blockers: Metadata review

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

agents/codecov-patch-coverage-planning-agent · trust review · confidence 67%

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

agents/material-ui-repository-contributor-agent · trust review · confidence 67%

Decision confidence

Decision confidence scan · balanced

2/16 results are low-confidence and need review before adoption.

high 2medium 12low 2

Codecov Patch Coverage Planning Agent

Address Metadata review, Package integrity before broader rollout.

medium

54/100

Missing: Metadata reviewMissing: Package integrity

agents/codecov-patch-coverage-planning-agent · trust review

Freshness distribution

Mostly fresh with a few aging entries

Median age 54 days; 11 fresh, 1 aging or stale of 12 scanned.

median 54d

Aging

91–180 days

8%

1 entry

Stale

> 180 days

0%

0 entries

Theme distribution

Themes are broadly spread across this view

55 distinct themes with no dominant one. Most common: evals, testing, regression.

Diverse

55 distinct themes across 16 scanned

Build repeatable eval suites that catch quality regressions in AI agent behavior before merge or release.

Level:advancedType:generalVerified:draft
Safety ✓ Privacy ✓

Expert Claude Code troubleshooting triage capability pack for diagnosing install failures, auth errors, MCP issues, sandbox blocks, and update regressions with source-backed triage matrices and privacy-safe support output.

Level:expertType:capability-packVerified:validated
Safety ✓ Privacy ✓

Slash command that reviews a pull request diff for security regressions: authentication and authorization gaps, injection surfaces, secret exposure, unsafe deserialization, and dependency risk introduced by the change.

Invocation:/pr-security-review [pr-number]
Safety ✓ Privacy ✓

Source-backed agent for turning Codecov patch coverage, project coverage, flags, components, carryforward behavior, PR comments, and changed-file context into targeted regression test plans.

PostToolUse hook that catches unintended UI changes by pixel-diffing a just-saved screenshot against its baseline with odiff.

Trigger:PostToolUse
Safety ✓ Privacy ✓
DeepEval logo
DeepEvalby Confident AI · submitted by oktofeesh1

Open-source Python framework for unit-testing LLM applications, agents, RAG pipelines, metrics, regression suites, and traces.

Expert reg-suit review skill for evaluating rendered UI image baselines, thresholds, snapshot storage, report artifacts, and visual QA release readiness.

Level:expertType:capability-packVerified:validated
Safety ✓ Privacy ✓
TruLens logo
TruLensby TruEra / Snowflake · submitted by oktofeesh1

Open-source evaluation and tracing framework for measuring AI agents, RAG systems, LLM apps, retrieval quality, feedback metrics, and trace-level regressions.

Slash command runbook for designing and running prompt evaluations: define tasks, success criteria, golden outputs, regression checks, and privacy-safe reporting using Anthropic test-and-evaluate guidance.

Invocation:/prompt-eval-runbook <feature-or-prompt-name>
Safety ✓ Privacy ✓
OpenAI logo

Source-backed guide for converting OpenAI Agents SDK traces into regression eval cases, trace grades, tool-call assertions, and release checks for agentic workflows.

A source-backed collection for building repeatable LLM eval and prompt testing workflows with open-source tools: prompt regression tests, RAG and agent metrics, human review datasets, traces, prompt optimization, and release gates.

OpenAI Evalsby OpenAI · submitted by JSONbored

Open-source framework from OpenAI for evaluating LLM and agent behavior with reusable eval definitions, grading logic, datasets, and regression workflows.

Source-backed Claude agent prompt for contributing to the official mui/material-ui monorepo using its AGENTS.md guidance, pnpm workspace filters, package build and test commands, component conventions, public error-message rules, API docs generation, visual regression and accessibility checks, and pre-PR checklist.

Source-backed agent for reviewing rendered frontend changes with screenshots, visual comparison evidence, viewport layout checks, keyboard/focus paths, accessibility scans, CLS risk, and privacy-safe QA artifacts.

User-created Claude Code custom slash command recipe for planning deeper tests around a file or function, including edge cases, property-style invariants, regression cases, and mutation-score gaps where the project already supports those tools.

Invocation:Create .claude/commands/test-advanced.md, then run: /test-advanced <file-or-function>
Safety ✓ Privacy ✓

Validate MCP server tools with contract-style tests to catch schema drift, unsafe behavior, and integration regressions early.

Level:advancedType:generalVerified:draft
Safety ✓ Privacy ✓