Source profileQuality 80/100Review permissions

affaan-m/ECC/.agents/skills/eval-harness/SKILL.md

eval-harness

Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles

Source repository stars
234,327
Declared platforms
1
Static risk flags
1
Last source update
2026-07-27
Source checked
2026-07-28

Decision brief

What it does—and where it fits

A formal evaluation framework for Claude Code sessions, implementing eval-driven development (EDD) principles.

Best for

    Not for

    • Tasks that require unconfirmed production actions or broad system permissions.
    • Environments where the pinned source and install steps cannot be inspected.

    Compatibility matrix

    Platform support, with evidence labels

    PlatformStatusEvidenceWhat to check
    CodexNot declaredNo explicit evidencePortability before use
    Claude CodeDeclaredSource recordInstall path and trigger
    CursorNot declaredNo explicit evidencePortability before use
    Gemini CLINot declaredNo explicit evidencePortability before use
    Open the compatibility checker

    Installation

    Inspect first. Install second.

    The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.

    Source-detected install commandSource
    npx skills add https://github.com/affaan-m/ECC --skill ".agents/skills/eval-harness"
    Safe inspection promptEditorial

    Inspect the Agent Skill "eval-harness" from https://github.com/affaan-m/ECC/blob/4e973d3eaf92d97f8d2e2d8abb39d8bdc8711b38/.agents/skills/eval-harness/SKILL.md at commit 4e973d3eaf92d97f8d2e2d8abb39d8bdc8711b38. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.

    Workflow

    What the source asks the agent to do

    1. 01

      Eval Workflow

      Review the “Eval Workflow” section in the pinned source before continuing.

      Review and apply the “Eval Workflow” source section.
    2. 02

      Pre-Implementation

      Creates eval definition file at .claude/evals/feature-name.md

      Creates eval definition file at .claude/evals/feature-name.md
    3. 03

      During Implementation

      Runs current evals and reports status

      Runs current evals and reports status
    4. 04

      Post-Implementation

      Generates full eval report

      Generates full eval report
    5. 05

      Phase 1: Define (10 min)

      Capability Evals: - [ ] User can register with email/password - [ ] User can login with valid credentials - [ ] Invalid credentials rejected with proper error - [ ] Sessions persist across page reloads - [ ] Logout clears session

      [ ] User can register with email/password[ ] User can login with valid credentials[ ] Invalid credentials rejected with proper error

    Permission review

    Static risk signals and limitations

    Runs scripts

    medium · line 56

    The documentation asks the agent to run terminal commands or scripts.

    npm test -- --testPathPattern="auth" && echo "PASS" || echo "FAIL"

    Runs scripts

    medium · line 59

    The documentation asks the agent to run terminal commands or scripts.

    npm run build && echo "PASS" || echo "FAIL"

    Evidence record

    Why each signal appears

    EvidenceSourceComputedTestedEditorial
    SignalValueEvidence typeMeaning
    Quality score80/100ComputedDocumentation, specificity, maintenance, and trust rules
    Repository stars234,327SourceRepository attention, not individual Skill quality
    Compatibility1 platformsSourceDeclared in the catalog source record
    Usage guideautomated source guideEditorialGenerated or reviewed according to the visible evidence level

    Pinned source

    Provenance and original SKILL.md

    Repository
    affaan-m/ECC
    Skill path
    .agents/skills/eval-harness/SKILL.md
    Commit
    4e973d3eaf92d97f8d2e2d8abb39d8bdc8711b38
    License
    MIT
    Collected
    2026-07-28
    Default branch
    main
    View the original SKILL.md

    Eval Harness Skill

    A formal evaluation framework for Claude Code sessions, implementing eval-driven development (EDD) principles.

    When to Activate

    • Setting up eval-driven development (EDD) for AI-assisted workflows
    • Defining pass/fail criteria for Claude Code task completion
    • Measuring agent reliability with pass@k metrics
    • Creating regression test suites for prompt or agent changes
    • Benchmarking agent performance across model versions

    Philosophy

    Eval-Driven Development treats evals as the "unit tests of AI development":

    • Define expected behavior BEFORE implementation
    • Run evals continuously during development
    • Track regressions with each change
    • Use pass@k metrics for reliability measurement

    Eval Types

    Capability Evals

    Test if Claude can do something it couldn't before:

    [CAPABILITY EVAL: feature-name]
    Task: Description of what Claude should accomplish
    Success Criteria:
      - [ ] Criterion 1
      - [ ] Criterion 2
      - [ ] Criterion 3
    Expected Output: Description of expected result
    

    Regression Evals

    Ensure changes don't break existing functionality:

    [REGRESSION EVAL: feature-name]
    Baseline: SHA or checkpoint name
    Tests:
      - existing-test-1: PASS/FAIL
      - existing-test-2: PASS/FAIL
      - existing-test-3: PASS/FAIL
    Result: X/Y passed (previously Y/Y)
    

    Grader Types

    1. Code-Based Grader

    Deterministic checks using code:

    # Check if file contains expected pattern
    grep -q "export function handleAuth" src/auth.ts && echo "PASS" || echo "FAIL"
    
    # Check if tests pass
    npm test -- --testPathPattern="auth" && echo "PASS" || echo "FAIL"
    
    # Check if build succeeds
    npm run build && echo "PASS" || echo "FAIL"
    

    2. Model-Based Grader

    Use Claude to evaluate open-ended outputs:

    [MODEL GRADER PROMPT]
    Evaluate the following code change:
    1. Does it solve the stated problem?
    2. Is it well-structured?
    3. Are edge cases handled?
    4. Is error handling appropriate?
    
    Score: 1-5 (1=poor, 5=excellent)
    Reasoning: [explanation]
    

    3. Human Grader

    Flag for manual review:

    [HUMAN REVIEW REQUIRED]
    Change: Description of what changed
    Reason: Why human review is needed
    Risk Level: LOW/MEDIUM/HIGH
    

    Metrics

    pass@k

    "At least one success in k attempts"

    • pass@1: First attempt success rate
    • pass@3: Success within 3 attempts
    • Typical target: pass@3 > 90%

    pass^k

    "All k trials succeed"

    • Higher bar for reliability
    • pass^3: 3 consecutive successes
    • Use for critical paths

    Eval Workflow

    1. Define (Before Coding)

    ## EVAL DEFINITION: feature-xyz
    
    ### Capability Evals
    1. Can create new user account
    2. Can validate email format
    3. Can hash password securely
    
    ### Regression Evals
    1. Existing login still works
    2. Session management unchanged
    3. Logout flow intact
    
    ### Success Metrics
    - pass@3 > 90% for capability evals
    - pass^3 = 100% for regression evals
    

    2. Implement

    Write code to pass the defined evals.

    3. Evaluate

    # Run capability evals
    [Run each capability eval, record PASS/FAIL]
    
    # Run regression evals
    npm test -- --testPathPattern="existing"
    
    # Generate report
    

    4. Report

    EVAL REPORT: feature-xyz
    ========================
    
    Capability Evals:
      create-user:     PASS (pass@1)
      validate-email:  PASS (pass@2)
      hash-password:   PASS (pass@1)
      Overall:         3/3 passed
    
    Regression Evals:
      login-flow:      PASS
      session-mgmt:    PASS
      logout-flow:     PASS
      Overall:         3/3 passed
    
    Metrics:
      pass@1: 67% (2/3)
      pass@3: 100% (3/3)
    
    Status: READY FOR REVIEW
    

    Integration Patterns

    Pre-Implementation

    /eval define feature-name
    

    Creates eval definition file at .claude/evals/feature-name.md

    During Implementation

    /eval check feature-name
    

    Runs current evals and reports status

    Post-Implementation

    /eval report feature-name
    

    Generates full eval report

    Eval Storage

    Store evals in project:

    .claude/
      evals/
        feature-xyz.md      # Eval definition
        feature-xyz.log     # Eval run history
        baseline.json       # Regression baselines
    

    Best Practices

    1. Define evals BEFORE coding - Forces clear thinking about success criteria
    2. Run evals frequently - Catch regressions early
    3. Track pass@k over time - Monitor reliability trends
    4. Use code graders when possible - Deterministic > probabilistic
    5. Human review for security - Never fully automate security checks
    6. Keep evals fast - Slow evals don't get run
    7. Version evals with code - Evals are first-class artifacts

    Example: Adding Authentication

    ## EVAL: add-authentication
    
    ### Phase 1: Define (10 min)
    Capability Evals:
    - [ ] User can register with email/password
    - [ ] User can login with valid credentials
    - [ ] Invalid credentials rejected with proper error
    - [ ] Sessions persist across page reloads
    - [ ] Logout clears session
    
    Regression Evals:
    - [ ] Public routes still accessible
    - [ ] API responses unchanged
    - [ ] Database schema compatible
    
    ### Phase 2: Implement (varies)
    [Write code]
    
    ### Phase 3: Evaluate
    Run: /eval check add-authentication
    
    ### Phase 4: Report
    EVAL REPORT: add-authentication
    ==============================
    Capability: 5/5 passed (pass@3: 100%)
    Regression: 3/3 passed (pass^3: 100%)
    Status: SHIP IT
    

    Alternatives

    Compare before choosing