Which AI Skill wins?
We put AI Skills head-to-head on the same tasks, generated by the same model, and judge the results blind. Pick a category to see who wins, and which skill fits your job.
How the arena works

Same task, same model
Every skill builds the identical set of tasks with one fixed base model, so results isolate the skill, not the model.

Blind judging
Outputs are rendered, anonymized, and scored 1 to 10 by independent judges on a rubric of axes chosen for that category.

Ranked + best-for
We average the scores into a leaderboard and a per-task winner, and note what each skill is genuinely best for.
Why benchmark AI Skills?
Skill marketplaces now list thousands of AI Skills, but there is no easy way to know which ones actually improve an agent’s output, or whether they help at all. Most public benchmarks rank models, not the skills layered on top of them.
The Skill Arena holds the model constant and varies only the skill, so every score reflects the skill’s real contribution. The goal is a practical, evidence-based answer to one question: which skill should I actually use for this job?
A note on method. Each category is judged on its own rubric of category-specific axes, and results come from a single run on one base model with a small task set, scored by AI judges rather than crowd votes. Treat them as directional evidence; open a category for its exact rubric, and re-test on your own workload before committing.
Methodology: skills build an identical set of tasks on one fixed model; every output is rendered, anonymized, and scored 1 to 10 by two independent blind judges on a category-specific rubric. Frontend-design results v1.0 · last updated July 2026.
AI Skills glossary
- AI Skill
- A reusable instruction file (SKILL.md) plus optional assets that teaches an AI agent how to do a specific job: design a frontend, build a slide deck, or analyze a spreadsheet.
- SKILL.md
- The single markdown file that defines a skill. It is read into the model on each turn, so its size directly drives token cost.
- Baseline (control)
- The same base model with no skill attached; the reference point for measuring how much a skill actually adds.
- Context tokens
- The total tokens processed per run, including the SKILL.md re-read into context each turn. A large skill file dominates this number.
- Blind judging
- Scoring anonymized outputs (labeled A to F) so judges cannot tell which skill produced which result.
Skill Arena FAQ
What is the Skill Arena?+
The Skill Arena is a blind-judged benchmark from Happycapy that pits AI Skills against each other on identical tasks. Every skill runs on the same base model, so the results reflect the skill's contribution, not the model's. Each category has its own leaderboard, side-by-side outputs, and a recommendation for what each skill is best at.
How are AI skills tested and scored?+
Each skill builds the same set of tasks with one fixed base model. Every output is rendered, anonymized, and scored 1-10 by two independent blind judges on a rubric of category-specific axes (for the frontend-design arena: visual polish, layout, typography, task completeness, and distinctiveness), then averaged across judges and tasks.
Which skill categories does the Skill Arena cover?+
Five arenas are live: frontend design, slides & PPTX, docs & writing, data & spreadsheets, and AI image generation, each comparing 6 skills across 3 tasks. Video is coming soon.
What has the Skill Arena found so far?+
In frontend design, three skills tied for the top overall score (8.0/10): frontend-design, design-taste-frontend and superdesign, with superdesign winning the landing-page task. In slides & PPTX, frontend-slides won overall (8.17) ahead of reveal.js. The takeaway: the best skill depends on the task, not one universal winner.
What is an AI Skill (SKILL.md)?+
An AI Skill is a reusable instruction file (SKILL.md) plus optional assets that teach an AI agent how to do a specific job, such as designing a frontend or building a slide deck. It is loaded into the model's context to steer its output, and its size affects how many tokens each run costs.
How often is the Skill Arena updated?+
Each category is versioned and dated, and we re-run tests and add new categories over time. The current frontend-design results are v1.0 (July 2026); the model, task set and date are listed on the category page so you can judge freshness.
Is the Skill Arena affiliated with the skill authors?+
No. Skills are independently selected from public, open directories and their source repos, then tested blind. We are not affiliated with the skill authors; every skill is judged blind on the same tasks and model.
How do I use the skills in the Skill Arena?+
Every skill tested is openly available from its source repo, and on Happycapy you can install and run any of them in your browser with no local setup. Open a category to see each skill's source, scores, and live demo.
Build with the skill that wins
Happycapy is an agent-native computer where you can run any of these skills right in your browser, with no setup. Browse the Skill Store for 2M+ skills, or jump straight in and start building.
