Skills Unit Testing

After writing a Skill, how can you confirm it actually works as expected? Manual testing is time-consuming and error-prone.

This article introduces how to build systematic test cases for a Skill, from script unit testing to end-to-end trigger testing.


Two Levels of Skill Testing

LevelTest TargetMethod
Script Unit TestingPython/JS scripts under scripts/pytest / standard unittest
Trigger-based End-to-End TestingWhether the Skill is triggered correctly and whether the output meets expectationsskill-creator's eval framework

Level 1: Script Unit Testing

Scripts are the most deterministic part of a Skill and are well suited to coverage with a standard unit testing framework.

Example

# File path: scripts/tests/test_clean_data.py
import pytest
import os
import tempfile
import csv

# Module under test
import sys
sys.path.insert(0, os.path.join(os.path.dirname(__file__), ".."))
from clean_data import remove_duplicates, fill_nulls, strip_whitespace

# ── Test: Deduplication ──────────────────────────────────────────
def test_remove_duplicates_basic():
    """Normal case: duplicate rows should be removed"""
    rows = [
        {"name": "example", "score": "90"},
        {"name": "example", "score": "90"},   # Duplicate row
        {"name": "EXAMPLE", "score": "80"},
    ]
    result = remove_duplicates(rows)
    assert len(result) == 2, "1 duplicate row should be removed"

def test_remove_duplicates_empty():
    """Edge case: empty list should not raise an error"""
    result = remove_duplicates([])
    assert result == []

# ── Test: Null Value Filling ──────────────────────────────────────
def test_fill_nulls_numeric():
    """Null values in numeric columns should be filled with 0"""
    rows = [{"value": "10"}, {"value": ""}, {"value": "20"}]
    result = fill_nulls(rows, col="value", fill_with="0")
    assert result[1]["value"] == "0"

# ── Test: Trimming Spaces ──────────────────────────────────────
def test_strip_whitespace():
    """Leading and trailing spaces in strings should be removed"""
    rows = [{"name": "  example  "}, {"name": "EXAMPLE"}]
    result = strip_whitespace(rows, col="name")
    assert result[0]["name"] == "example"
    assert result[1]["name"] == "EXAMPLE"

# ── Integration Test: Reading a Real File ──────────────────────────────
def test_process_real_file():
    """Create a temporary CSV file and test the complete processing flow"""
    with tempfile.NamedTemporaryFile(mode="w", suffix=".csv",
                                    delete=False, newline="") as f:
        writer = csv.DictWriter(f, fieldnames=["name", "score"])
        writer.writeheader()
        writer.writerow({"name": "  example  ", "score": "90"})
        writer.writerow({"name": "  example  ", "score": "90"})  # Duplicate
        writer.writerow({"name": "EXAMPLE",     "score": ""})    # Null value
        tmp_path = f.name

    try:
        from clean_data import process_file
        result = process_file(tmp_path)
        assert result["removed_rows"] == 1
        assert result["fixed_values"] == 1
    finally:
        os.unlink(tmp_path)

Run the tests:

# 进入 Skill 目录后运行所有测试
cd my-skill/
pytest scripts/tests/ -v

# 生成覆盖率报告
pytest scripts/tests/ --cov=scripts --cov-report=term-missing

Output:

scripts/tests/test_clean_data.py::test_remove_duplicates_basic  PASSED
scripts/tests/test_clean_data.py::test_remove_duplicates_empty  PASSED
scripts/tests/test_clean_data.py::test_fill_nulls_numeric       PASSED
scripts/tests/test_clean_data.py::test_strip_whitespace         PASSED
scripts/tests/test_clean_data.py::test_process_real_file        PASSED

5 passed in 0.12s

Level 2: Trigger-based End-to-End Testing

Trigger testing verifies whether Claude will use the Skill when a user makes a request.

Test cases are defined in JSON format, containing user prompts and expected behavior assertions.

Example

[
  {
    "id": "trigger_basic",
    "prompt": "Help me analyze this sales data CSV, find monthly trends and generate a statistical summary",
    "assertions": [
      {
        "type": "skill_triggered",
        "skill": "csv-analyzer",
        "description": "Complex analysis requests should trigger csv-analyzer"
      }
    ]
  },
  {
    "id": "trigger_with_file",
    "prompt": "I uploaded example_sales.csv, please analyze the data distribution of each column",
    "assertions": [
      {
        "type": "skill_triggered",
        "skill": "csv-analyzer"
      },
      {
        "type": "output_contains",
        "keyword": "Statistics",
        "description": "The output should contain statistical information"
      }
    ]
  },
  {
    "id": "no_trigger_simple",
    "prompt": "What format is CSV?",
    "assertions": [
      {
        "type": "skill_not_triggered",
        "skill": "csv-analyzer",
        "description": "Simple knowledge questions should not trigger the Skill"
      }
    ]
  }
]

Run trigger tests (requires the eval tool provided by skill-creator):

Example

# Run the test suite and generate a visual report
python -m scripts.run_eval \
  --eval-set evals/trigger-eval.json \
  --skill-path csv-analyzer/ \
  --model claude-sonnet-4-20250514

# Generate an HTML report for manual review
python eval-viewer/generate_review.py \
  --results evals/results/ \
  --output evals/review.html

Test case prompts should be sufficiently complex. For simple requests like "read CSV", even if triggering fails, it doesn't mean the Skill has a problem, because Claude can handle it on its own. Good test cases should be multi-step requests users would actually make, with clear output requirements.


Test Case Coverage

TypeScenarios to cover
Positive triggeringComplex tasks, containing keywords, clear output format requirements
Negative triggeringSimple Q&A, only mentioning relevant terms but not needing the Skill
Boundary inputsEmpty files, oversized files, files with unsupported formats
Error recoveryWhether the error message is clear when the file path is wrong
Other extensions