Skip to main content

Browse the directory

Showing 10 resources for "regression"
Saved
Active

Source-backed filter active — add entries to compare trust side by side.

Trust snapshot

10 results in this view

Claimed
0%(0/10)

2 trust signals differ in this sample: Source provenance, Submitter

Signals differ on Source provenance, Submitter — add entries to compare before you install.

Rollout signal scan

2 rollout risk signals in current results

Biggest gaps: metadata review, package integrity. 1 entries have 2+ required gaps.

10 scanned

Install payload

Install payload is mixed and needs spot-checking.

watch

70% (7/10)

Adoption queue

Browse adoption queue · balanced

3/10 visible results are in hold tier and need mitigation before adoption.

ready 0caution 7hold 3
caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

skills/claude-code-troubleshooting-triage-capability-pack · trust review · confidence 67%

Codecov Patch Coverage Planning Agent

1 blockers: Metadata review

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

agents/codecov-patch-coverage-planning-agent · trust review · confidence 67%

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

agents/material-ui-repository-contributor-agent · trust review · confidence 67%

OpenAI Evals

1 blockers: Metadata review

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

tools/openai-evals · trust review · confidence 67%

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

skills/reg-suit-visual-regression-review-capability-pack · trust review · confidence 67%

caution

50/100

Request metadata review from maintainers or internal owners.

Collect package checksum or signed artifact information.

hooks/screenshot-visual-regression · trust review · confidence 67%

DeepEval

2 blockers: Metadata review, Install payload

hold

36/100

Request metadata review from maintainers or internal owners.

Add install/config payload for reproducible team rollout.

Collect package checksum or signed artifact information.

tools/deepeval · trust review · confidence 50%

Decision confidence

Decision confidence scan · balanced

3/10 results are low-confidence and need review before adoption.

high 0medium 7low 3

Codecov Patch Coverage Planning Agent

Address Metadata review, Package integrity before broader rollout.

medium

54/100

Missing: Metadata reviewMissing: Package integrity

agents/codecov-patch-coverage-planning-agent · trust review

OpenAI Evals

Address Metadata review, Package integrity before broader rollout.

medium

54/100

Missing: Metadata reviewMissing: Package integrity

tools/openai-evals · trust review

DeepEval

Hold adoption until Metadata review, Package integrity are resolved.

low

36/100

Missing: Metadata reviewMissing: Package integrityMissing: Install payload

tools/deepeval · trust review

Freshness distribution

Mostly fresh with a few aging entries

Median age 54 days; 9 fresh, 1 aging or stale of 10 scanned.

median 54d

Aging

91–180 days

10%

1 entry

Stale

> 180 days

0%

0 entries

Oldest entries in this view

Theme distribution

Themes are broadly spread across this view

36 distinct themes with no dominant one. Most common: testing, capability-pack, evaluation.

Diverse

36 distinct themes across 10 scanned

Expert Claude Code troubleshooting triage capability pack for diagnosing install failures, auth errors, MCP issues, sandbox blocks, and update regressions with source-backed triage matrices and privacy-safe support output.

Level:expertType:capability-packVerified:validated
Safety ✓ Privacy ✓

Slash command that reviews a pull request diff for security regressions: authentication and authorization gaps, injection surfaces, secret exposure, unsafe deserialization, and dependency risk introduced by the change.

Invocation:/pr-security-review [pr-number]
Safety ✓ Privacy ✓

Source-backed agent for turning Codecov patch coverage, project coverage, flags, components, carryforward behavior, PR comments, and changed-file context into targeted regression test plans.

PostToolUse hook that catches unintended UI changes by pixel-diffing a just-saved screenshot against its baseline with odiff.

Trigger:PostToolUse
Safety ✓ Privacy ✓
DeepEval logo
DeepEvalby Confident AI · submitted by oktofeesh1

Open-source Python framework for unit-testing LLM applications, agents, RAG pipelines, metrics, regression suites, and traces.

Expert reg-suit review skill for evaluating rendered UI image baselines, thresholds, snapshot storage, report artifacts, and visual QA release readiness.

Level:expertType:capability-packVerified:validated
Safety ✓ Privacy ✓
TruLens logo
TruLensby TruEra / Snowflake · submitted by oktofeesh1

Open-source evaluation and tracing framework for measuring AI agents, RAG systems, LLM apps, retrieval quality, feedback metrics, and trace-level regressions.

OpenAI Evalsby OpenAI · submitted by JSONbored

Open-source framework from OpenAI for evaluating LLM and agent behavior with reusable eval definitions, grading logic, datasets, and regression workflows.

Source-backed Claude agent prompt for contributing to the official mui/material-ui monorepo using its AGENTS.md guidance, pnpm workspace filters, package build and test commands, component conventions, public error-message rules, API docs generation, visual regression and accessibility checks, and pre-PR checklist.

Promptfoo logo

Open-source prompt testing and red-teaming framework for LLM outputs, regressions, evaluations, and security checks.

Safety · Privacy ✓