AI Image Generation skills
Skills that turn a brief into a finished image — posters, infographics and storyboards.
We gave six general-purpose image-generation skills the same three briefs, and rendered every result with the same fixed image model, OpenAI gpt-image-2. Each skill only shapes the prompt and workflow; because the image model is held constant, the differences reflect the skill, not the model.
Leaderboard
JimLiu (Baoyu) · The most reliable all-rounder — its disciplined prompt workflow won the storyboard and topped the arena overall.
ConardLi · Fast starts from a matching template across many image types.
AgriciDaniel · Balanced, on-brief results across image types with sensible defaults.
jezweb · Teaching disciplined, well-structured prompts for photographic scenes.
smixs (Serge Shima) · Art-directed, on-brief visuals — it won the poster and tied for best on the storyboard.
inference.sh · Clean, correct, information-dense results — it won the infographic and is by far the leanest (cheapest) skill here.
Overall = mean of the per-task overall scores from 2 blind judges.
Every result, side by side
Every skill built the same brief on the same model. Tap any result to open the live page.
Game onboarding screen
Winner: baoyu-image-genA polished game character-onboarding screen (16:9): a large hero portrait, class intro, icon-based ability cards, a difficulty indicator, starter stats and a primary action button — immersive fantasy game UI with layered panels, cinematic lighting and strong hierarchy, not a website landing page.
Best overall entry — the large left-anchored hero portrait bleeds beautifully into a structured right panel, all text elements are present and legible, difficulty gems read clearly, and the START ADVENTURE button commands attention with strong cinematic lighting throughout.
The most premium, game-ready layout — full widescreen use with a left lore panel, dominant centered hero name, a clean horizontal abilities row with icon cards, and a right-side difficulty indicator; the only weakness is that ability card body text is rendered at a scale too small to read comfortably.
Strong adherence with all required elements present and clearly labeled, though the layout feels slightly cramped with the character portrait pushed left and the UI panel dominating centrally rather than framing the hero dramatically.
The right-column ability card layout is reasonably clean but the stats are oddly buried in a single bottom bar rather than given a dedicated panel, and the character name block on the lower left competes visually with the portrait rather than anchoring it.
Strong character art and atmosphere, but the ability cards are listed as plain text rows rather than distinct icon-based cards, and the overall layout leans more toward a vertical list UI than a cinematic game screen.
The fiery cinematic atmosphere is the strongest of the set and the horizontal ability card row is well-proportioned, but small descriptive flavor text beneath the cards is unreadably tiny and the difficulty indicator dots are barely discernible.
Fashion campaign poster
Winner: ai-image-generatorA vertical 4:5 DTC eyewear social-media ad: a confident young model in stylish sunglasses (the hero product) against a vibrant, non-sky colour backdrop, bright and energetic, with graphic negative space carrying a 'SEE THE DAY DIFFERENTLY' headline and a 'NEW SUNGLASSES COLLECTION' subtitle.
Tight portrait crop, vibrant electric-blue backdrop, bold white headline stacked cleanly at top with subtitle directly below — textbook DTC eyewear execution with sharp lens detail and excellent negative space hierarchy.
Mirror-lens sunglasses pop brilliantly against the vivid blue, centered composition with strong negative space at top and bottom delivers a premium yet energetic DTC ad — arguably the most polished and shippable execution.
Strong low-angle masculine energy, vivid blue studio backdrop, clean white type, good negative space at top — slightly less refined subtitle placement at the very bottom edge, but overall a compelling, social-media-ready execution.
Bold cobalt-blue wardrobe blending into backdrop is a creative DTC energy choice, model is natural and warm, but the color-on-color model/background merge slightly reduces product-focus clarity and compositional separation.
Blue backdrop reads almost sky-like (horizon gradient, outdoor light quality) which partially violates the non-sky studio requirement, and the subtitle is pushed to an awkward bottom corner with weak typographic hierarchy.
Strong blue color energy and good low-angle model framing, but headline and subtitle are crammed together at the top in a narrow single line with little breathing room, and the background again reads outdoor/sky rather than solid studio.
Editorial infographic
Winner: gpt-image-2 (garden)An editorial magazine-style infographic about coffee (vertical 4:5): bold cropped photography, colour blocks, oversized numbers, a pull quote and a simple chart in a dynamic asymmetrical lifestyle-magazine layout — energetic but readable, not vintage or watercolour.
Excellent asymmetric editorial composition with bold full-bleed coffee photography, all four stats stacked vertically with strong typographic hierarchy, clear pull quote, and a well-labelled bar chart showing numeric values (12.0 kg, 9.1 kg, 8.2 kg) for all three countries — minor deduction for the very small source/footnote text at bottom reducing overall chart legibility.
Strongest entry — all four stats, title, pull quote, and a clearly labelled three-country bar chart with numerical values are all present and correct; the marble-texture photo, tight horizontal stat strip, and restrained dark palette deliver a genuinely contemporary editorial magazine aesthetic with excellent hierarchy and legibility.
All required elements present and legible, with a dynamic diagonal arrow graphic adding editorial energy; the bar chart uses percentage-scale bars with all three countries labelled, though the bars are quite short and the chart section feels compressed compared to the strong upper half.
All required content present and correctly spelled including the bar chart with three labelled countries, bold editorial look with strong photo crop, though the bar chart values are very small and the layout leans slightly symmetrical rather than truly dynamic.
All required content is present and correctly spelled, the latte-art photo is warm and high quality, but the layout is nearly identical to B in structure — symmetrical two-column stat grid, flat orange blocks, and the bar chart bars for all three countries appear almost equal in length which undermines the data differentiation.
All required content is present and correctly spelled, orange/black palette is bold and magazine-like, but the bar chart bars appear nearly identical in length with minimal visual differentiation between the three countries, and the layout feels somewhat rigid and symmetric.
How we scored ai image generation
Two independent blind judges scored every output 1 to 10 on each of these axes, then we averaged across judges and tasks. These axes are specific to ai image generation, while other arenas use their own rubric.
- Prompt adherence
- Whether every element the brief asked for is actually present.
- Visual quality
- Craft and finish, realism, and freedom from glitches or artifacts.
- Composition
- Framing, balance, focal hierarchy and use of space.
- Aesthetics
- Color, lighting, mood and overall taste.
- Text & legibility
- Correct, legible, well-placed text — a real differentiator for image models.
Limitations. This arena is a single run on gpt-image-2 with 3 tasks, scored by AI judges rather than crowd votes. Treat the results as directional evidence, not a definitive ranking; re-test on your own workload before committing.
How each skill performed, and who it’s for
baoyu-image-gen
A power-user image workflow: write every prompt to a file first, with reference images, aspect ratios, batch and many providers.
Best for: The most reliable all-rounder — its disciplined prompt workflow won the storyboard and topped the arena overall.
Watch out: Ops-heavy: the prompt-file/preferences workflow is overkill for a quick one-off image.
gpt-image-2 (garden)
A gpt-image-2-native template library: 80+ structured templates across 18 categories (poster, infographic, storyboard and more).
Best for: Fast starts from a matching template across many image types.
Watch out: Template defaults can read generic; it trailed the workflow-driven skills on polish.
banana
A 'Creative Director' that interprets intent, picks a domain mode, and builds a 6-component prompt brief.
Best for: Balanced, on-brief results across image types with sensible defaults.
Watch out: Middle of the pack; its Gemini-tuned guidance is less differentiated once the model is fixed to gpt-image-2.
ai-image-generator
A structured 5-part narrative prompt framework (image type, subject, environment, technical specs, constraints) plus model routing.
Best for: Teaching disciplined, well-structured prompts for photographic scenes.
Watch out: Its rigid framework was least effective here on layout- and text-heavy briefs like the infographic.
visual-skills
A cinematic 'director' discipline: a model-specific GPT Image slot template, an anti-cliché banned-word list and exact-text quoting.
Best for: Art-directed, on-brief visuals — it won the poster and tied for best on the storyboard.
Watch out: Its reference-heavy reading workflow is demanding, and it was weaker on the dense infographic layout.
ai-image-generation
A multi-model catalogue router: pick the right model for the job, then write a natural-language prompt.
Best for: Clean, correct, information-dense results — it won the infographic and is by far the leanest (cheapest) skill here.
Watch out: Little prompt methodology of its own; its model-selection edge is neutralized when the model is fixed.
The tasks, head-to-head
How each skill scored on the three briefs, and how quality trades off against token cost.
| Skill | Game onboarding screen | Fashion campaign poster | Editorial infographic | Avg |
|---|---|---|---|---|
| baoyu-image-gen | 9.0 | 9.0 | 8.0 | 8.7 |
| gpt-image-2 (garden) | 8.0 | 8.5 | 9.0 | 8.5 |
| banana | 9.0 | 7.0 | 8.0 | 8.0 |
| ai-image-generator | 7.0 | 9.0 | 7.0 | 7.7 |
| visual-skills | 7.0 | 7.0 | 9.0 | 7.7 |
| ai-image-generation | 7.0 | 8.0 | 7.0 | 7.3 |
Cell = the skill’s score on that task (darker = higher). A ring marks the task winner.
Token cost & efficiency
Every skill ran on the same base model, so token use reflects the skill itself: how much it writes (generated) and how large its SKILL.md is (context processed). Summed across the 3 tasks; lower is cheaper.
Context includes the skill’s SKILL.md re-read into the model each turn, so a larger skill file drives up cost, while the no-skill baseline is cheapest.
What we learned
- baoyu-image-gen led the board overall (8.67/10), with gpt-image-2 (garden) second (8.5).
- ai-image-generator won the fashion campaign poster, gpt-image-2 (garden) won the editorial infographic, and baoyu-image-gen won the game onboarding screen — no single skill swept, the best skill depends on the job.
- Leaner held up: the smallest skill, ai-image-generation (~7K context tokens), still ranked #6 — a bigger prompt framework did not guarantee better images once the model was fixed.
- Text was decisive: skills that quoted the exact copy and planned where it sits scored highest on legibility, which separated the top skills.
- Because the image model is fixed to gpt-image-2, model-selection-heavy skills lose their main lever and compete purely on prompt craft.
Frequently asked questions
What is the best AI skill for generating images?+
In our blind test baoyu-image-gen scored highest overall (8.67/10), with gpt-image-2 (garden) a close second (8.5). Every skill rendered with the same fixed model (gpt-image-2), so the ranking reflects prompt and workflow craft, not the image model.
Which image skill is best for the different tasks?+
It varies by task: in our tests ai-image-generator won the fashion campaign poster, gpt-image-2 (garden) won the editorial infographic, and baoyu-image-gen won the game onboarding screen. Match the skill to the job rather than expecting one universal winner.
Does the image generation model or the skill matter more?+
We held the model constant (gpt-image-2) so we could measure the skill. A good skill clearly helps — by quoting exact text, planning layout and adding art direction — but the gains are about prompt craft, and skills whose main value is picking a model lose that advantage when the model is fixed.
How do these skills handle text inside images?+
Text legibility was the biggest differentiator. gpt-image-2 renders text well, and the skills that quoted the exact copy and planned where it goes produced correct, legible headlines, labels and captions; weaker prompts led to garbled or missing text.
How is the Image arena scored, and is it fair?+
Every skill builds the same three briefs and renders on the same fixed model (gpt-image-2), so results isolate the skill. Each image is anonymized (A to F) and scored 1 to 10 by two independent blind judges on prompt adherence, visual quality, composition, aesthetics and text legibility, then averaged across judges and tasks.
Where can I find and install these image skills?+
All are openly available: baoyu-image-gen (JimLiu), ai-image-generation (inference.sh), visual-skills (smixs), banana (AgriciDaniel), gpt-image-2 (ConardLi) and ai-image-generator (jezweb). On Happycapy you can install any of them and run them in your browser with no setup.
Build with these skills, no setup required
Every skill in this ai image generation arena installs on Happycapy in one click and runs in your browser on the same model. Start free, or explore 2M+ skills in the Skill Store.


















