Evaluate whether the skill improves real engineering work.
1. Positive and negative cases
Test tasks that should activate the skill and similar tasks that should not. Measure both missed activation and irrelevant activation.
2. Outcome rubric
Score correctness, scope, verification and report clarity. A skill that produces longer answers without better work has not demonstrated value.
3. Regression fixtures
Keep representative repositories or small fixtures with known problems. Re-run them after instruction changes and compare the resulting edits and checks.
Worked scenario
A new trigger phrase improves activation for Compose bugs but starts activating on general Kotlin exercises. The negative set catches that regression.
Apply it
Create five intended tasks, five near misses and two adversarial cases with expected behavior.
Check your understanding
You can identify a skill revision that makes performance worse. Explain the decision and show evidence from your implementation or design. If you cannot demonstrate it yet, revisit the relevant section before continuing.