Skip to content
On this page

Reference · Working notes · Noord

Slop score: point to the problem

How we mark visual rough edges, rate their effect, and turn the review into useful corrections. The method behind Noord’s slop scoring tool.

Download this method as Markdown

“This looks like slop” is a reaction, not a useful brief. Which bit? What is wrong with it? What does it stop someone from seeing or doing? We built a small review tool to make ourselves answer those questions.

Slop score is our human review of avoidable visual roughness. We draw on the actual screenshots, explain the concerns, then rate the app as a whole. It is not an AI detector. People are perfectly capable of making a mess without assistance.

Point to it

We look for decisions that feel unconsidered and get in the way of the work. A familiar pattern is not automatically slop. Neither is a plain interface. The question is whether a choice serves this task, in this context.

Every mark needs a reason someone else can act on. “These three actions have equal weight, so we cannot tell which completes the review” gives us a job. “Bad vibes” gives us a meeting.

  • Spacing and alignment: unrelated things appear grouped, related things drift apart, or edges and baselines miss each other without a reason.
  • Weak hierarchy: the main decision disappears among equally loud cards, labels, or buttons.
  • Generic decoration: a flourish takes space and attention but does not help this particular screen.
  • Inconsistent components: the same action changes treatment, or a component ignores the shared system without a useful exception.
  • Unclear copy: a label hides what will happen, feedback is vague, or the tone does not belong in the situation.

Draw, then explain

The tool puts the screenshot beside its review controls. We switch between apps and views, draw a red box around a concern, choose a category and severity, and write a short note. Marks can be removed or undone. The image stays unchanged underneath.

For our design-model study, every app has six captures: desktop, open dialog, phone, phone dialog, dark desktop, and dark phone. The app-level rating stays unavailable until all six are marked reviewed and every box has a note. A screen with no concerns can still be reviewed. An unreviewed screen cannot quietly count as perfect.

Model names and earlier design scores are hidden in the review interface. That removes a cue; it does not erase our memory of work we have already seen. Reviews save to the local project, with an export and import option. The JSON keeps the screenshot references, marks, notes, and rating together so the review can travel with the work.

How we weigh it

After inspecting the views, we choose one overall rating from 0 to 5. Multiply it by 20 for a slop score from 0 to 100. Lower is better. The scale is deliberately coarse: we are choosing a level of roughness, not pretending to measure taste to two decimal places.

We weigh the consequence, how often the problem recurs, and how much of the task it affects. One obscured primary action can matter more than several slightly awkward gaps. Seeing the same defect in three captures is evidence of a recurring problem, not three separate penalties.

The categories organize the review; they have no percentage weights. Severity helps us decide what to repair first; it is not added into a formula. Box count and covered area never determine the score. Otherwise the fastest route to a better result would be drawing smaller rectangles. Very efficient. Entirely useless.

How we weigh it
RatingSlop scoreWhat we see
00No visible concerns in the inspected states
120A few small, isolated rough edges
240Noticeable cleanup in several places
360Recurring problems compete with the content
480Careless or generic treatment dominates the design
5100The visual treatment makes the interface hard to use throughout

Fix what matters

Each mark uses one of three severity labels. Use them to order the corrections, and explain the effect in the note. A serious usability or accessibility failure remains a blocker even when the rest of the screen looks good.

  • 1 · Cosmetic distraction
  • 2 · Makes the design harder to read
  • 3 · Obscures important content or controls

No magic average

We keep slop score separate from the original design score and the technical checks. A design can have a strong direction and still need cleanup. A beautifully consistent interface can still solve the wrong problem. Mixing those into one total would hide the useful difference.

No completed review means no score—not zero. Zero means no visible concerns in the states we actually inspected. It does not mean the product has passed every test. The rating is a reviewer’s judgment; compare versions with the same task, captures, and scale, and record disagreements rather than smoothing them away.

Where it belongs

We place this review after a working iteration exists and before we ask a partner to judge it. First check that the app runs and the important states are reachable. Then use the screenshots to inspect the visual decisions. Turn the findings into small corrections, make another pass, and review the new version.

Keep the original capture and score. A repaired version gets its own record. We want to see whether the correction helped, not gradually erase the evidence of what needed fixing.

  1. 01CaptureBrowser + saved versionThe same states, sizes, and themes
  2. 02ReviewPerson + slop toolAnnotated concerns and one overall rating
  3. 03RepairDesigner or bounded agent taskA new version with specific corrections
  4. 04Try againPerson + working appChecked fixes, then partner review
If a task needs a new decision, return it to the planning conversation before continuing.

Use it, too

A screenshot cannot tell us how long a button takes to respond, whether focus gets lost, or how the screen feels on a moving train with one hand occupied. Try the interaction, read the words in context, and use the actual device. Slop review sits beside that work; it does not replace it.

An agent can capture a repeatable set of states or implement an approved correction. The judgment about which choice helps the person—and whether the work is ready to share—stays with people.

Our first pass

We built the tool while reviewing nine apps made by ChatGPT, Claude, and Grok. The first export contains 19 marked areas across 5 apps. These are our actual boxes, shown on the original captures. The overall ratings are still blank; this is an annotation pass, not a completed slop-score comparison.

19marked areas across 5 apps

Slop score · not rated yet0–100 · lower is better

Tools, in order

Here is where each tool helps. Use an equivalent you already work with if it fits the task and your project’s access requirements.

  1. Before marking

    Browser and accessibility checksA versioned project folder

    Check the app runs, then save matching views with the version, viewport sizes, and themes. Keep private project data out of shared captures.

    Carry forward An unchanged screenshot set tied to one version.

  2. During review

    Noord slop review tool

    Inspect every view, mark specific concerns, and explain their effect. Choose one overall rating after the image review; keep it separate from the design score and technical checks.

    Carry forward A saved review with boxes, notes, severity, and a completed rating or an explicit blank.

  3. After review

    NotionClaude, ChatGPT, or GrokBrowser and accessibility checks

    Agree on the corrections in Notion. Give a designer or agent a bounded task, then inspect the changed app in context and save a new review without overwriting the original.

    Carry forward Checked corrections and a decision about partner review.

Review the work with partners

We make and revise the work with Claude, ChatGPT, or Grok, then put selected versions on our Studio project pages for partners to review. Check the preview before sharing it. You can do the same with a private prototype or shared document; use a workspace that can access the files you need.

Name the version, say what changed, and ask the question you need answered. Keep feedback with that version and discuss conflicting requests before making the next changes. Confirm approval separately, and keep confidential work in a restricted space.

A question for this review

Which marked problem makes this task hardest, and did the next version actually remove it?

Delegate the routine work

Capture the agreed states or implement one approved correction with a before-and-after check.

Deciding what counts as a problem, assigning the overall score, and approving the work remain human judgments.

Product documentation

Follow these links for the providers’ current instructions. They describe capabilities, not independent evidence for Noord’s model preferences.