Production Quality Assessment
CORPUS is being built in the open. The assessment described here is built and being integrated into scoring. Expect details to evolve.
The Production Quality Assessment feeds the quality dimension of scoring. Because its results will affect what a contribution earns, this page lays out what it measures, against what, and what it deliberately refuses to do.
Why there is no taste score
There is no universal scale of good production. A great punk record and a great pop master are close to technical opposites: one raw, narrow, and distorted on purpose, the other wide, clean, and sculpted for dynamics. A crushed, maximally loud master is a defect on an audiophile jazz record and the whole point of a club track. Production quality is entangled with genre, era, instrumentation, and intent; a single number for it would put a number on taste while actually measuring proximity to whatever the model happened to be fed.
So the system asks two smaller questions that have real answers:
- Is the audio technically intact? Bandwidth-limited, mangled by a lossy codec, clipped, broken? This is objective and independent of style.
- Is the track unusual for what it is trying to be? Does it sit far outside the range that well-made records of its own kind normally occupy?
The second question is anomaly detection, and the output is a production-risk estimate: a map of where to look, with the measurements that earned each flag. Risk asks a person to go and check; it is not a verdict.
The reference library
Style-conditioned analysis needs a trustworthy picture of what well-made records of a given kind sound like. The base of the system is a curated library of around 600 studio recordings, sorted into roughly 30 broad style buckets (Pop, Jazz, Metal, Hip-Hop, Classical, Reggae, Afrobeats, and so on).
- Curated, not scraped. Every track earned its slot, and the reasons are written down: who produced, engineered, and mastered it, why it gets cited as a reference, and how confident the entry is on each point. Many entries trace back to the recordings that working mastering engineers repeatedly name as references in the trade press and professional forums, the records the people who master for a living keep on the desk.
- Anti-references included. A handful of deliberately loud, flat records document a real and legitimate way of making music; a system that takes context seriously has to hold those too.
- Three excerpts per track. The loudest stretch from each third of a song, so a quiet intro or one stray peak never speaks for the whole record. Around 600 tracks become roughly 1,800 reference excerpts.
What gets measured
Every feature is a classical, readable audio descriptor, the kind a mastering engineer would recognize and could argue with. Two separate sets do two separate jobs:
- Production families describe how the record is made: loudness (integrated LUFS, loudness range, true peak), dynamics (RMS, peak, crest factor, peak-to-loudness ratio), spectral balance across six bands, source quality (roll-off, high-frequency energy, effective bandwidth, codec-cutoff detection), and stereo image (width, balance, mono compatibility, bass in the sides).
- Context features describe what kind of music the track is: tempo, rhythmic density, structure, harmonic content, timbral fingerprint. These never touch the quality result; their only job is to pick the right reference neighborhood before any comparison happens.
No trained black-box scorer sits in this chain. If the tool flags a problem, it can say which measurable thing looks wrong, in words a person can check against their own ears.
How the scoring works
For each production family the system computes a robust z-score: how far the track sits from the median of its reference neighborhood, in units of that neighborhood's normal spread. Three rules keep the measurement fair:
- Style sets the bar, except for integrity. Loudness, dynamics, spectral balance, and stereo are judged only against the matched style pool; a punk track is never held to a jazz record's dynamics. Source quality is always judged against the entire library, because whether audio is intact has nothing to do with style.
- Defects are flagged, differences are left alone. A track brighter than its references is ignored (a defensible choice); a suspiciously dull one is flagged (a missing top end usually means a lossy or band-limited source).
- Some defects have a hard floor. A genuine 8 kHz downsample or a 64 kbps MP3 carries a penalty no friendly reference pool can offset. The floor only drops on evidence; lossless files are never penalized on suspicion.
What the assessment reports
The assessment returns a picture, not a verdict: the overall risk, per-family subscores, the specific measurements that moved the result (each next to the reference median it was compared against), the nearest reference tracks, which style pool was chosen and with what confidence, and plain-language warnings when an objective floor kicked in. When the system cannot tell which style a hybrid track belongs to, it says so.
The system can tell the difference between a track that is badly made and a track that is simply not to someone's liking. How its findings translate into points, and what of them contributors get to see, is still being settled as part of the scoring integration.
Next: Your Dashboard.