COSP10 Research Hub All articles
Machine Learning Engineering

Training on Garbage: Why Scaling LLMs Without Cleaning Data Is a Billion-Dollar Mistake

COSP10 Research Hub
Training on Garbage: Why Scaling LLMs Without Cleaning Data Is a Billion-Dollar Mistake

Photo: data quality machine learning training noisy dataset visualization, via blog.allegro.tech

The Scaling Assumption That Quietly Fails

For several years, the dominant narrative in large language model development has been straightforward: more parameters, more compute, better results. The scaling laws published by researchers at OpenAI and DeepMind seemed to confirm this intuition mathematically. Feed a model more data, expand its architecture, and performance climbs predictably.

The problem is that those scaling laws carry an implicit assumption that most practitioners never examine closely enough — that the data being scaled is of consistent, representative quality. Strip away that assumption, and the entire framework begins to unravel.

At COSP10, we think of this in terms of signal fidelity. Every training corpus is, at its core, a transmission medium. The question is not how much data you are sending through the model, but how much of that data carries genuine signal versus noise that degrades the learning process. When the signal-to-noise ratio of a training dataset deteriorates, scaling becomes an amplifier of problems rather than a solution to them.

What Data Noise Actually Looks Like in Practice

Data noise in LLM training is rarely as obvious as corrupted files or garbled text. The more insidious forms arrive in shapes that pass automated validation checks without difficulty.

Consider near-duplicate content. Web-scraped corpora often contain thousands of near-identical product descriptions, templated news articles, and syndicated blog posts. On the surface, these look like legitimate training examples. In practice, they cause the model to overweight certain phrasings and syntactic patterns, effectively narrowing the distribution of language the model internalizes. Researchers at the Allen Institute for AI demonstrated in their work on data deduplication that removing near-duplicates from training sets produced measurable downstream improvements in model coherence — without adding a single additional parameter.

There is also the problem of domain contamination. When benchmark test sets leak into training corpora — a phenomenon now documented across several major datasets — models achieve artificially inflated evaluation scores that evaporate the moment they encounter genuinely novel prompts. This is not a hypothetical concern. Multiple organizations have published post-mortems describing how their models appeared to hit performance milestones that could not be reproduced in production environments, a pattern our earlier coverage of lab-to-live performance erosion explored in depth.

A third category, perhaps the most underappreciated, is temporal noise. Training corpora assembled from web crawls contain information that was accurate at one point in time but has since become incorrect, superseded, or contextually misleading. When a model is trained on contradictory factual claims about the same subject — accurate in 2019, inaccurate by 2023 — it does not neatly resolve the conflict. It learns to be confidently inconsistent.

Case Studies in Wasted Compute

The financial consequences of training on degraded data are difficult to overstate. A single large-scale training run for a frontier-class model can consume millions of dollars in GPU-hours on platforms like AWS, Google Cloud, or Azure. When that run is built on a compromised corpus, those resources are not merely inefficient — they produce a model that requires additional remediation work, fine-tuning cycles, or in some cases, complete retraining.

One documented example comes from the open-source community. The early releases of several multilingual models suffered from a phenomenon researchers termed "language bleed," where low-resource language data was so sparse and inconsistently formatted that the model learned to hallucinate plausible-sounding but grammatically incoherent text in those languages. The root cause was not model architecture. It was a failure to audit the quality distribution of the multilingual training set before the run began.

In enterprise settings, similar failures manifest differently but stem from the same root. Organizations that fine-tune base models on proprietary document repositories frequently discover that their internal data is far noisier than anticipated. Legal documents contain boilerplate that crowds out substantive content. Customer service logs are riddled with formatting artifacts. Technical manuals include version-specific instructions that contradict one another across document generations. Without pre-training data audits, these issues compound during fine-tuning and produce models that confidently generate incorrect internal-facing outputs.

A Diagnostic Framework Before the Training Clock Starts

The most cost-effective intervention in LLM training is one that happens before a single GPU warms up. The following framework, grounded in signal-processing principles, offers a structured approach to data quality assessment.

Step 1: Distribution Mapping. Before training begins, generate statistical profiles of your corpus across key dimensions: token frequency distributions, sentence length histograms, domain source proportions, and temporal spread. Anomalies in any of these dimensions warrant investigation. A corpus where 40 percent of tokens originate from a single domain is not a general-purpose training set, regardless of its raw size.

Step 2: Deduplication at Multiple Granularities. Exact-match deduplication is necessary but insufficient. Implement fuzzy deduplication using MinHash or SimHash algorithms at the document level and, where computationally feasible, at the paragraph level. The goal is to ensure that the model encounters genuine variation in its training signal rather than repeated reinforcement of narrow patterns.

Step 3: Benchmark Contamination Auditing. Cross-reference your training corpus against the test splits of every evaluation benchmark you intend to use. This is a non-negotiable step. Tools such as the contamination detection utilities released alongside several recent model papers can automate much of this process.

Step 4: Temporal Coherence Scoring. For corpora assembled from web crawls, tag documents with their source timestamps and assess whether contradictory factual claims can be traced to temporal drift. Where possible, prefer recency-weighted sampling for factual domains and explicitly exclude outdated content.

Step 5: Human Spot-Auditing. Automated quality metrics will not catch everything. Randomly sample several hundred documents from your corpus and have domain-knowledgeable reviewers assess them for coherence, factual plausibility, and formatting integrity. This step is unglamorous, but it consistently surfaces categories of noise that automated pipelines miss.

The Economics of Getting This Right

There is a persistent organizational tendency to treat data preparation as a cost center and model training as the value-generating activity. This framing inverts the actual economics. A week of rigorous data auditing that prevents a compromised training run pays for itself many times over, particularly as compute costs at scale remain substantial.

The signal-processing metaphor that informs COSP10's editorial approach is instructive here. In communications engineering, no amount of amplification can recover a signal that was degraded at the source. The same principle applies to language model training. Scale amplifies what is already present in the data. If what is present is noise, scale produces a very large, very expensive noise amplifier.

Organizations that internalize this principle — treating data quality as a first-order engineering constraint rather than an afterthought — will consistently outperform those that chase parameter counts. The smarter model is rarely the bigger model. It is the one trained on the cleaner signal.

All Articles

Related Articles

From Lab to Live: Understanding Why Production Environments Erode Machine Learning Model Performance

From Lab to Live: Understanding Why Production Environments Erode Machine Learning Model Performance

When Accuracy Becomes a Mirage: Diagnosing False Confidence in Machine Learning Systems

When Accuracy Becomes a Mirage: Diagnosing False Confidence in Machine Learning Systems

The Noise Floor Problem: How Flawed Data Pipelines Are Quietly Undermining AI Model Performance

The Noise Floor Problem: How Flawed Data Pipelines Are Quietly Undermining AI Model Performance