Stop Claude From Shipping Ugly Landing Pages (3 Fixes)

Abstract tech illustration: Stop Claude From Shipping Ugly Landing Pages (3 Fixes)

You've shipped a landing page from Claude Code and felt embarrassed to send the link. The code is clean, the components are correct, the accessibility is fine — but something about it screams "generated." It's not your prompt. It's that Claude has never been shown, on purpose, what good actually looks like.

Anthropic knows this. Their own official web-artifacts-builder skill explicitly warns developers to avoid the "AI slop" look — excessive centered layouts, purple gradients, uniform rounded corners, and the Inter font (Artifacts Builder skill docs). If the team that trained the model is calling out the failure mode by name, you can stop blaming yourself.

Below: the five tells I look for on every AI-generated page, why the model produces them, and three fixes I run on real client work, ranked by effort.

The 5 tells that give away an AI-generated landing page

Once you see these, you can't unsee them. This is my own working checklist from shipping client pages — not from a study — but every item maps to a design pattern that recurs across generated output.

Tell 1 — spacing that's technically correct but emotionally flat. Everything is 16px, 24px, 32px. No breathing room around the hero, no tension between sections. Real designers use a spacing scale that jumps — 8, 16, 24, 48, 96 — so the eye knows where to rest.

Tell 2 — no elevation hierarchy. Every card has the same shadow, the same border, the same weight. Nothing feels primary. On a real product page, one thing is loud, three things are medium, everything else is quiet.

Tell 3 — default radii. 8px rounded corners on everything: buttons, cards, images, inputs. Pro design uses radius as a language — pills for actions, sharp edges for data, soft rectangles for content.

Tell 4 — gradient overuse. Purple to pink on the hero, blue to teal on the CTA, some indigo shimmer in the footer. There's a well-known explanation for the purple: many AI coding tools default to Tailwind's indigo-500, and that default propagated through training data into model output (DEV Community writeup).

Tell 5 — no type scale. Headline bold, body regular, done. No display weight, no tracking adjustment on the H1, no size ratio between H2 and body. Everything reads at the same volume.

Why Claude defaults to this look

Language models converge on a visual formula because they're trained on far more generic SaaS landing pages than distinctive, art-directed designs (Open Intelligence / Medium analysis). When you say "make it modern," the statistical average of "modern" is the middle of that distribution.

It gets worse under the hood. A recent cross-environment study found AI-generated webpages had a 68% rate of at least one browser/device compatibility issue, versus 40% for human-authored pages (arXiv paper). So the output isn't just visually average — it's fragile in ways you won't see until a prospect opens it on Safari.

You're not going to prompt your way out of averageness with "make it beautiful." You have to show Claude what good looks like, on purpose. That's what the next three fixes do.

Two notes on tooling before the fixes:

  • Claude Artifacts — the panel that renders generated HTML/React landing pages — is available on every plan tier including Free, Pro, Max, Team, and Enterprise (Albato).
  • Anthropic Labs released Claude Design, a dedicated visual prototyping tool for full websites and landing pages, as a research preview on April 17, 2026, running on Claude Opus 4.7 at launch (Clauder Navi). Useful, but the fixes below work in plain Claude Code and don't depend on it.

Fix 1 — Reference injection (lowest effort, highest return)

Instead of "build a pricing section," say "build a pricing section, and here is the design system I want you to match." Then paste three things:

  1. A screenshot of a benchmark site. I use Linear, Vercel, and Stripe because their published components are ruthless about spacing and hierarchy.
  2. The actual token values you want. Color palette in hex, font stack, spacing scale, radius scale — as literal numbers, not adjectives.
  3. One negative example. A screenshot of an AI-generated page with the instruction: do not produce anything that looks like this.

The third piece is what most people skip and it's the highest leverage. Claude follows negative constraints better than positive ones.

A working prompt scaffold:

Build a 3-tier pricing section.

DESIGN TOKENS (use these literally):
- Colors: bg #0A0A0A, surface #141414, text #FAFAFA, muted #8A8A8A, accent #E85D3C
- Font: "Inter Display" for headings, "Inter" for body
- Spacing scale (non-linear): 4, 8, 16, 24, 48, 96
- Radius: 4px cards, 999px CTAs, 2px inputs
- Shadows: only the primary CTA gets elevation. Everything else is flat.
- No gradients. Anywhere.

REFERENCES (attached):
- linear-pricing.png — match this spacing + hierarchy
- stripe-pricing.png — match this type scale
- ai-slop-example.png — DO NOT produce anything resembling this

CONSTRAINTS:
- Type scale ratio: H1 = 3.5rem/72px, H2 = 2rem/32px, body = 1rem/16px
- One tier is "primary" — visibly heavier via elevation, not via a gradient

In practice, rebuilding a pricing section this way collapses hours of Tailwind hand-tuning into a short prompt-and-review pass. Your mileage varies with how tight your token spec is, but the pattern holds across every client page I've shipped since.

Fix 2 — The critic loop (consistent quality across every page)

Reference injection gets you one good output. A critic loop gets you consistent good output across every page you ship.

The architecture:

  • Builder agent — regular Claude Code writing the component.
  • Critic agent — a separate Claude instance whose only job is to score the output against a rubric.
  • Rubric — the 5 tells above, plus brand specifics: correct type scale, primary CTA has distinct elevation, spacing scale is non-linear, no unauthorized gradients.
  • Loop — if the critic score is below 8/10, the builder rewrites with the critic's feedback attached. Repeat until pass or max_iter hits.

Sketch:

MAX_ITER = 6
THRESHOLD = 8

component = builder.generate(spec, tokens, references)

for i in range(MAX_ITER):
    review = critic.score(component, rubric)
    if review.score >= THRESHOLD:
        break
    component = builder.revise(component, review.violations)

deploy(component, meta={"critic_score": review.score, "iterations": i + 1})

The critic prompt is where your taste lives. Skeleton:

You are a senior product designer reviewing generated landing-page code.
Score 1-10 against this rubric. Return JSON: {score, violations[]}.

RUBRIC (fail any one = max score 6):
1. Spacing scale is non-linear (e.g., 8/16/24/48/96), not uniform 16/24/32.
2. Elevation hierarchy exists: exactly one primary element has distinct shadow.
3. Radius is used semantically (pills for actions, rectangles for content).
4. Zero gradients unless explicitly listed in tokens.
5. Type scale has a real ratio: H1 >= 2x body, tracking adjusted on H1.
6. Color palette matches provided hex values exactly.

For each violation, cite the exact CSS class or line.

The critic is honestly the more important agent — it has your taste encoded. A weak critic means a weak loop, no matter how good the builder is. If you can't write the rule, the critic can't enforce it. Spend an afternoon writing out what you actually mean by good. Not vibes. Rules. Sizes allowed. Colors allowed. Shadow values allowed.

I run this on a home server in WSL Ubuntu so it can chew through five or six iterations without me babysitting. On Claude Sonnet 5.5 — released September 28, 2026 at $2/M input and $10/M output tokens, with output more than 30% faster than Sonnet 5 (Anthropic) — a six-iteration loop on a single component is cheap enough to run on every commit.

Fix 3 — The style recipe library (repeatable at agency scale)

Every time you land on a component pattern that works — a hero, a pricing table, a feature grid, a testimonial row — save three files:

  • The final prompt that produced it.
  • The reference screenshots you used.
  • The token bundle (JSON).

All three live in a folder on your machine, organized by component type:

~/recipes/
  hero/
    minimal-dark/
      prompt.md
      references/
        linear-hero.png
        vercel-hero.png
        slop-example.png
      tokens.json
  pricing/
    three-tier-card/
      prompt.md
      references/
      tokens.json
  feature-grid/
    ...

New client project starts, you don't start from scratch. Call the recipe, swap tokens.json for the client's brand, run the critic loop. What used to be a multi-hour design pass becomes a short pipeline. On agency landing pages, this is the difference between a two-day turnaround and a two-week one.

Honest limits

A few things this stack still doesn't fix:

  • Custom illustration. Claude does not draw. You need a designer or a separate image model.
  • Photography direction. The model will pick generic stock unless you provide URLs.
  • Motion and interaction feel. You can generate the animation code, but the timing curves usually need a human pass.
  • Cross-browser fragility remains real per the compatibility study above — always render-test on Safari and mobile Chrome before shipping.

Anthropic's own guidance on avoiding the AI-slop look ships as a developer-invoked skill, not something applied automatically to every generation. Which is why fixes 1–3 are on you.

Where bizflowai.io fits in

bizflowai.io is practical AI automation for solopreneurs and small teams — and the landing pages that sell those systems get the same builder + critic + recipe treatment described above. The critic rubric is the durable asset. Once a client's brand is encoded as literal rules — allowed hex values, allowed spacing steps, allowed shadow tokens — every new page, email template, and pitch deck runs through the same loop. That's how a one-person shop ships pages that don't look one-person-shop.


Want more like this?

I publish practical AI automation, GenAI engineering, and faceless content workflows — tutorials, teardowns, and working code, not theory.

Planning an AI automation project or need a second opinion on your architecture? Reach out through the site — I read every message.

Visit bizflowai.io for services, case studies, and AI consulting from Lazar Milićević / BizFlowAI.

Frequently asked questions

How do I spot an AI-generated landing page?

Look for five tells: flat linear spacing (everything at 16/24/32px), no elevation hierarchy so nothing feels primary, default 8px rounded corners on every element, overused gradients like purple-to-pink hero and blue-to-teal CTA, and no type scale where headline is just bold and body is regular. Once you notice these patterns, you can identify AI output in under two seconds.

What is reference injection in AI design prompting?

Reference injection is a prompting technique where you paste three things alongside your build request: a screenshot of a benchmark site (like Linear, Vercel, or Stripe), the actual design tokens you want (hex colors, font stack, spacing and radius scales), and a negative example — an AI-generated screenshot with instructions not to produce anything like it. Negative constraints are the highest-leverage piece.

How does a critic loop improve AI-generated design?

A critic loop uses two agents: a builder that writes the component and a separate critic agent that scores the output against a rubric covering spacing, elevation, type scale, and gradient rules. The critic returns a score from 1-10 and specific violations. If the score is below 8, the builder rewrites using the feedback. This loops until it passes or hits a max iteration count.

When should I use reference injection vs a critic loop?

Use reference injection when you need one good output quickly — it is low effort and high return, useful for single sections or one-off prompts. Use a critic loop when you need consistent quality across every page you ship. Reference injection gets you one good result; the critic loop encodes your taste into a repeatable system that produces reliable output at scale.