# Slop score: point to the problem

How we mark visual rough edges, rate their effect, and turn the review into useful corrections. The method behind Noord’s slop scoring tool.

**TL;DR**

Point to the rough edges, explain what they get in the way of, then rate the whole app. Lower is better; more red boxes does not mean more slop. The rectangles are there to help us fix the work, not win an argument about taste.

“This looks like slop” is a reaction, not a useful brief. Which bit? What is wrong with it? What does it stop someone from seeing or doing? We built a small review tool to make ourselves answer those questions.

Slop score is our human review of avoidable visual roughness. We draw on the actual screenshots, explain the concerns, then rate the app as a whole. It is not an AI detector. People are perfectly capable of making a mess without assistance.

## Point to it

We look for decisions that feel unconsidered and get in the way of the work. A familiar pattern is not automatically slop. Neither is a plain interface. The question is whether a choice serves this task, in this context.

Every mark needs a reason someone else can act on. “These three actions have equal weight, so we cannot tell which completes the review” gives us a job. “Bad vibes” gives us a meeting.

- Spacing and alignment: unrelated things appear grouped, related things drift apart, or edges and baselines miss each other without a reason.

- Weak hierarchy: the main decision disappears among equally loud cards, labels, or buttons.

- Generic decoration: a flourish takes space and attention but does not help this particular screen.

- Inconsistent components: the same action changes treatment, or a component ignores the shared system without a useful exception.

- Unclear copy: a label hides what will happen, feedback is vague, or the tone does not belong in the situation.

## Draw, then explain

The tool puts the screenshot beside its review controls. We switch between apps and views, draw a red box around a concern, choose a category and severity, and write a short note. Marks can be removed or undone. The image stays unchanged underneath.

For our design-model study, every app has six captures: desktop, open dialog, phone, phone dialog, dark desktop, and dark phone. The app-level rating stays unavailable until all six are marked reviewed and every box has a note. A screen with no concerns can still be reviewed. An unreviewed screen cannot quietly count as perfect.

Model names and earlier design scores are hidden in the review interface. That removes a cue; it does not erase our memory of work we have already seen. Reviews save to the local project, with an export and import option. The JSON keeps the screenshot references, marks, notes, and rating together so the review can travel with the work.

![Noord’s slop review tool: screenshot views, annotation controls, and a separate rating panel](https://noord.dev/polder/benchmarks/design-models-study/theme-recheck/tool/slop-review.jpg)

Our review tool keeps each mark beside the image it refers to. Model names and earlier scores stay hidden.

## How we weigh it

After inspecting the views, we choose one overall rating from 0 to 5. Multiply it by 20 for a slop score from 0 to 100. Lower is better. The scale is deliberately coarse: we are choosing a level of roughness, not pretending to measure taste to two decimal places.

We weigh the consequence, how often the problem recurs, and how much of the task it affects. One obscured primary action can matter more than several slightly awkward gaps. Seeing the same defect in three captures is evidence of a recurring problem, not three separate penalties.

The categories organize the review; they have no percentage weights. Severity helps us decide what to repair first; it is not added into a formula. Box count and covered area never determine the score. Otherwise the fastest route to a better result would be drawing smaller rectangles. Very efficient. Entirely useless.

| Rating | Slop score | What we see |
| --- | --- | --- |
| 0 | 0 | No visible concerns in the inspected states |
| 1 | 20 | A few small, isolated rough edges |
| 2 | 40 | Noticeable cleanup in several places |
| 3 | 60 | Recurring problems compete with the content |
| 4 | 80 | Careless or generic treatment dominates the design |
| 5 | 100 | The visual treatment makes the interface hard to use throughout |

### Fix what matters

Each mark uses one of three severity labels. Use them to order the corrections, and explain the effect in the note. A serious usability or accessibility failure remains a blocker even when the rest of the screen looks good.

- 1 · Cosmetic distraction

- 2 · Makes the design harder to read

- 3 · Obscures important content or controls

### No magic average

We keep slop score separate from the original design score and the technical checks. A design can have a strong direction and still need cleanup. A beautifully consistent interface can still solve the wrong problem. Mixing those into one total would hide the useful difference.

No completed review means no score—not zero. Zero means no visible concerns in the states we actually inspected. It does not mean the product has passed every test. The rating is a reviewer’s judgment; compare versions with the same task, captures, and scale, and record disagreements rather than smoothing them away.

## Where it belongs

We place this review after a working iteration exists and before we ask a partner to judge it. First check that the app runs and the important states are reachable. Then use the screenshots to inspect the visual decisions. Turn the findings into small corrections, make another pass, and review the new version.

Keep the original capture and score. A repaired version gets its own record. We want to see whether the correction helped, not gradually erase the evidence of what needed fixing.

1. Capture — Browser + saved version
   Output: The same states, sizes, and themes

2. Review — Person + slop tool
   Output: Annotated concerns and one overall rating

3. Repair — Designer or bounded agent task
   Output: A new version with specific corrections

4. Try again — Person + working app
   Output: Checked fixes, then partner review

### Use it, too

A screenshot cannot tell us how long a button takes to respond, whether focus gets lost, or how the screen feels on a moving train with one hand occupied. Try the interaction, read the words in context, and use the actual device. Slop review sits beside that work; it does not replace it.

An agent can capture a repeatable set of states or implement an approved correction. The judgment about which choice helps the person—and whether the work is ready to share—stays with people.

[The release quality check](https://noord.dev/polder/qa-gate)

[Turn findings into bounded tasks](https://noord.dev/polder/delegating-with-prompts)

## Our first pass

We built the tool while reviewing nine apps made by ChatGPT, Claude, and Grok. The first export contains 19 marked areas across 5 apps. These are our actual boxes, shown on the original captures. The overall ratings are still blank; this is an annotation pass, not a completed slop-score comparison.

19 marked areas across 5 apps. Scores: not rated yet. Red boxes show the reviewer’s original selections; their number and area do not determine a score.

[Inspect the red-box previews](https://noord.dev/polder/choosing-design-models#slop-score) · [Download annotation coordinates and source hashes](https://noord.dev/polder/benchmarks/design-models-study/slop-annotations.json)

[Read the design-model study](https://noord.dev/polder/choosing-design-models)

### Product documentation

- [Design-model study and review method](https://noord.dev/polder/choosing-design-models#the-test)

- [Screenshot annotations and source hashes](https://noord.dev/polder/benchmarks/design-models-study/slop-annotations.json)

Product documentation describes capabilities, not independent evidence for Noord’s preferences.

Source: https://noord.dev/polder/slop-score

### Tools, in order

#### Before marking

[Browser and accessibility checks](https://www.w3.org/WAI/test-evaluate/preliminary/) · A versioned project folder

Check the app runs, then save matching views with the version, viewport sizes, and themes. Keep private project data out of shared captures.

Carry forward: An unchanged screenshot set tied to one version.

#### During review

[Noord slop review tool](/polder/slop-score#the-tool)

Inspect every view, mark specific concerns, and explain their effect. Choose one overall rating after the image review; keep it separate from the design score and technical checks.

Carry forward: A saved review with boxes, notes, severity, and a completed rating or an explicit blank.

#### After review

[Notion](https://www.notion.com/) · [Claude](https://claude.ai/), [ChatGPT](https://chatgpt.com/), or [Grok](https://grok.com/) · [Browser and accessibility checks](https://www.w3.org/WAI/test-evaluate/preliminary/)

Agree on the corrections in Notion. Give a designer or agent a bounded task, then inspect the changed app in context and save a new review without overwriting the original.

Carry forward: Checked corrections and a decision about partner review.

#### Review the work with partners

We make and revise the work with Claude, ChatGPT, or Grok, then put selected versions on our Studio project pages for partners to review. Check the preview before sharing it. You can do the same with a private prototype or shared document; use a workspace that can access the files you need.

Name the version, say what changed, and ask the question you need answered. Keep feedback with that version and discuss conflicting requests before making the next changes. Confirm approval separately, and keep confidential work in a restricted space.

For this method: Which marked problem makes this task hardest, and did the next version actually remove it?

#### Delegate the routine work

Capture the agreed states or implement one approved correction with a before-and-after check.

Keep with a person: Deciding what counts as a problem, assigning the overall score, and approving the work remain human judgments.

[Sequence prompts, then delegate](https://noord.dev/polder/delegating-with-prompts)

[Choose a model for the design task](https://noord.dev/polder/choosing-design-models)