Study Guide

Coffee Quality Grading: Counting Defects vs Scoring Cups

Separate defect counting from sensory scoring, classify primary and secondary defects, and build repeatable cup scores for the Coffee Quality Grading Exam.

Updated September 202610 min readStudy GuideWineConquer
Simon Kelly

Simon Kelly

WineConquer Editorial Team

The core difficulty in coffee quality grading is not that any single task is obscure; it is that the work alternates between two incompatible judgment modes. Green grading is observational and arithmetic: you sort a weighed sample, assign each defect to a category, and convert counts. Cupping is comparative perception: you rate attributes of the liquid against an internal scale. Errors happen when the modes bleed into each other — a bad defect count editing palate scores, or an averaged cup score hiding a non-uniform lot. This article trains the separation itself: sequence the tasks, keep the records apart, and practice verification and calibration until each mode stands on its own evidence.

Defect counting and cup scoring are two different judgment modes

Treating green-grading arithmetic and sensory evaluation as one blended task corrupts both. Keep them separate: classify and count defects on the sample first, then cup and score, and never let one record edit the other.

The two modes use different attention. Green grading asks: what is physically present in this weighed sample, and which named category does each defective bean belong to? Every conclusion is verifiable by inspection, cutting, or smell. Cupping asks: what does this liquid express in fragrance, flavor, acidity, body, and balance, relative to reference points in your memory? One mode produces counts; the other produces calibrated ratings. Confusing them means neither output is trustworthy.

The practical risk runs in both directions. If you cup a lot knowing its green sample showed heavy insect damage, expectation pressure can drag flavor and cleanliness scores below what the liquid actually shows. A visually clean sample can do the opposite and inflate scores. Sequence the work: complete and set aside the green grading sheet before cupping, and record cup impressions before comparing notes with anyone. Two sheets, filled independently, is the discipline to rehearse until automatic.

Primary versus secondary defects: why the categories are not interchangeable

Primary defects such as full black, full sour, insect damage, and fungus damage carry full weight; secondary defects such as broken beans, shells, hulls, and floaters carry lesser weight. Misassigning a category changes the computed grade.

Classification rests on observable cues. A full black bean is discolored through the whole bean and often smells moldy or rotten when cut. A full sour bean shows yellowish-brown discoloration and typically gives a fermented, vinegar-like odor on cutting. Insect damage announces itself with pinholes or bores, sometimes with dust inside. Partials, broken beans, shells, and floaters belong to the secondary side: the bean is defective in form or completeness but not necessarily spoiled. Each assignment should rest on a confirming observation, not a glance.

The judgment calls live at the edges. A bean can be both broken and discolored, or black only in patches. Rather than stacking multiple counts on one bean, the habit to build is: identify its dominant verified characteristic, assign one category per bean under the protocol you are studying, and set genuinely ambiguous beans aside with a note. Cut suspects open whenever the exterior is unclear — internal color and odor settle most black-versus-sour questions quickly. Drill on a tray of deliberately mixed beans until category assignment feels like sorting, not debating.

This is where a comparison table earns its place in your notes.

ObservationLikely categoryConfirming check
Dark discoloration through the whole beanPrimary (full black)Cut the bean; check internal darkness and musty odor
Yellow-brown tint, fermented smell when cutPrimary (full sour)Cut and smell; confirm sour odor inside
Pinholes, bores, or fine dust insidePrimary (insect damage)Look for the bore hole; check for frass
Broken or chipped but otherwise soundSecondary (broken)Confirm no internal discoloration
Boat-shaped hollow beanSecondary (shell)Compare shape against a reference bean
Black patches only, extent unclearDepends on protocol thresholdEstimate affected proportion; note as borderline

Worked scenario: a fragrant lot with a high defect count

In a simplified practice case, a lot smells floral and clean in the cup while the green sample shows several insect-damaged beans. The mistake is lowering flavor and clean-cup scores to match the count; the better decision is reporting each result on its own sheet.

Setup, with illustrative numbers: a grader inspects a weighed green sample, finds beans with clear borings, and computes a defect count that sits near the top of the tolerated range for the buyer's specification. The same day, the cupped lot is fragrant, sweet, and balanced. The plausible mistake: the grader, influenced by the ugly green sheet, writes down restrained flavor and cleanliness scores so the paperwork looks consistent. The count migrated into the palate record and contaminated it.

The better decision: the green grading sheet states the insect-damage count and category with the verification behind it, and the cupping sheet scores what the liquid actually expressed, attribute by attribute. Why it matters: an infestation problem at origin and a cup-quality result are different findings that trigger different responses — a pest-control conversation versus a pricing or blending decision. When the two records are allowed to disagree, the disagreement itself is diagnostic. Treat the numbers here as practice arithmetic, not as fixed tolerances from any particular standard.

Calibration: building scores you can repeat next week

A cup score is only useful if the same coffee would earn roughly the same score from you later. Calibration means anchoring each attribute to fixed reference points and checking your drift with blind repeat samples.

Start with anchors. Assemble a small reference set: one coffee you score mid-range for acidity, one low, one high, and equivalents for body or balance if you prefer. Re-cup these periodically and compare against your earlier notes. Keep vocabulary fixed — use a structured flavor-wheel approach so that citric, stone fruit, and fermented each mean the same thing in every session. Vocabulary drift is the quiet cause of score drift; if fermented means one thing in September and another in November, your uniformity scores are not comparable across time.

Then test repeatability directly. Cup the same coffee blind twice within a session, or across two days, and compare the sheets. Set your own practice tolerance band in advance — for example, individual attributes within one point of each other and a total score within a band you choose. If the repeat falls outside the band, check session conditions first: roast level, grind, water temperature, and resting time all move scores. Only after ruling those out should you conclude that your palate calibration drifted and needs re-anchoring. These bands are learning milestones for you, not predictions of any grading outcome.

Exercise: green-sample defect triage with a self-check rubric

Run a triage on a mixed green sample: weigh it, separate defects into primary and secondary piles, verify each bean by cutting or smelling, then recount. Score yourself on verification quality, not speed.

Procedure: weigh the sample and record the weight. Spread it on a dark tray and make a first pass pulling obvious defects into labeled piles. Make a second pass through smaller sub-samples for the beans you missed. Cut every suspect bean whose exterior is ambiguous, and log each removal by category and count. Deliberately mixed trays work well for this: ask a roaster for floor sweepings or defective fractions, or build your own by combining a clean lot with a handful of each defect type you can source.

Grade the exercise against this rubric: every removed bean carries a named category; every primary classification is backed by a confirming observation such as cut color, internal odor, or a bore hole; a full recount on a second day matches the first tally; and borderline beans appear in a separate noted pile instead of being forced into a category. If your recount differs from your first count, identify which category drifted — that category, not counting in general, is the specific skill your next drill should target. Repeat the exercise with a new mixed tray until the rubric passes cleanly twice in a row.

  • Every removed bean has a named primary or secondary category
  • Every primary call has a confirming observation on record
  • An independent recount matches the first tally
  • Ambiguous beans are logged separately, not forced into a pile

Scenario: pass, regrade, or hold a receiving lot

In a simplified receiving case, a lot's defect count is acceptable but several of five bowls show a fermented note and the cups disagree with each other. The better decision is to hold and re-cup with freshly drawn samples, not to average the scores upward.

Setup and mistake: five bowls are cupped from a lot. Three are clean and sweet; two carry a distinct fermented, winey edge. The receiving clerk is tempted to average across bowls, and the average clears the lot comfortably. That average is the mistake — it erases exactly the information multiple bowls exist to reveal. Non-uniformity within a lot is a finding in its own right, and averaging per-cup scores treats a mixed lot as if it were a consistent one.

The better decision: record each bowl separately, flag the uniformity failure alongside the sensory descriptor, and hold the lot. Re-cup using newly drawn samples from different bags in the lot to determine whether the fermented note is localized or widespread. Why it matters: a consistent lot that cups mediocre presents a pricing question; an inconsistent lot presents a process-control question at drying or fermentation, and accepting it on an average means the buyer absorbs risk nobody priced. Per-cup records make the two cases distinguishable; averages make them identical.

The decision pattern generalizes into a table worth keeping beside your cupping forms.

DecisionEvidence patternTypical next step
AcceptDefect count within specification and cups uniformProceed with receiving documentation
Hold / re-cupCups disagree with each other, or the count is borderlineDraw fresh samples from different bags and re-cup
Reject or renegotiatePrimary defects exceed tolerance, or a taint repeats across fresh samplesEscalate with both the green and cupping records
Investigate processA repeatable sensory fault traces to drying or fermentationFeed findings back to the supplier with specific descriptors

Clean protocols, honest records, and your readiness checklist

Grading conclusions are only as good as the samples and records behind them: rinse between cups, block fragrance carryover, write results before discussion, and never adjust a number to fit an expected outcome.

Hygiene in cupping is a calibration issue, not etiquette. Work in an unscented room, free of perfume, food aromas, and scented cleaning products. Rinse spoons and glasses thoroughly between samples; residual aromas from a previous cup attach themselves to the next one and masquerade as flavor notes. Keep water temperature and steeping routine identical across all bowls in a session, because an inconsistent procedure produces differences you will wrongly attribute to the coffees. Avoid strongly flavored food and drink beforehand for the same reason.

Ethics covers the record itself. Scores and defect counts are commercial documents: someone may price, accept, or reject a container on what you write. Record first impressions immediately, before anyone speaks, and leave submitted sheets unedited — corrections go on a new, dated sheet. Disclose any interest you hold in the lot being graded. Before you consider yourself ready, verify each item on this checklist honestly: it marks study milestones, not a predicted result.

  • You can classify a defect to a named category and verify it by cutting or smell without consulting a key
  • Your blind re-cup of a reference coffee falls inside your own tolerance band
  • You can write a pass, hold, or reject recommendation citing both the green and cupping records
  • You keep defect counts and sensory scores on separate, independently completed sheets

References and further reading

Use these references to explore the concepts and check the latest information from the relevant organizations.

Continue your preparation

FAQ

Frequently Asked Questions

Practical answers to help you apply the guidance for Coffee Quality Grading Exam.

Do I need to memorize exact numeric tolerances for defects and scores?
Learn the structure first — defect categories, the cupping attributes, and how per-cup scores are recorded — then attach the specific numbers from the protocol your exam references. Do not carry counts or conversion ratios from one grading system into another; they differ.
Is defect counting done on green or roasted coffee?
Green grading examines the green sample and produces a count by category. Sensory scoring examines the roasted, cupped liquid. The two results stay on separate records and answer different questions about the same lot.
How can I practice calibration with a limited coffee budget?
A small reference set of two or three coffees, re-cupped blind across sessions, supports real calibration work. For defect practice, one mixed tray of defective green beans can be sorted, recounted, and re-sorted many times before it needs replacing.
What if a single bean seems to fit two defect categories?
Apply the single-category rule from the protocol you are studying: classify by the bean's dominant verified characteristic and note the ambiguity. Cutting the bean open usually settles whether spoilage or physical damage dominates.
Do my self-check scores predict how I will perform on the exam?
No. The tolerance bands and rubric checks in this guide are learning milestones that tell you when your skills are internally consistent. Administrative details and any scoring rules for the exam itself come from the issuing body.

Keep Reading

Related Study Guides

Explore related guides and preparation topics.