---
title: "Production Dataset Preparation Pipeline: A Case Study"
author: "Acadify Engineering Team"
author_role: "AI & Software Engineering Team"
date: "October 01, 2026"
categories: [Case Studies]
description: "An empirical case study on building a production dataset preparation pipeline: solving deduplication, layout parsing, PII scrubbing, and chunking for AI."
---

# Production Dataset Preparation Pipeline: A Case Study

By **Acadify Engineering Team** (AI & Software Engineering Team) on October 01, 2026

## Executive Summary

In enterprise artificial intelligence implementations, dataset preparation is frequently underestimated as mundane data munging, yet it represents the single most deterministic factor in downstream model accuracy, safety, and operational reliability. This case study documents the end-to-end engineering of a production-grade dataset preparation pipeline designed to transform 2.4 million raw, heterogeneous organizational artifacts into structured, high-signal corpora for retrieval-augmented generation (RAG) and domain fine-tuning.

By shifting from naive programmatic extraction to a rigorous four-stage pipeline—structural normalization, Locality-Sensitive Hashing (LSH) deduplication, automated PII sanitization, and hierarchical layout-aware chunking—the engineering team eliminated 23.9% redundant volume, improved context recall from 61.2% to 84.7%, and eradicated sensitive data leakage without degrading domain context.

## The Challenge: The Hidden Costs of Naive Extraction

Enterprise data lakes contain thousands of multi-page PDF reports, slide decks, spreadsheet extracts, and markdown technical specs. Naive programmatic scrapers (such as standard unformatted text dumps) produce massive structural corruption:

- **Loss of Document Geometry:** Multi-column layouts, callout boxes, and headers/footers bleed into body prose, creating hallucination-inducing noise.

- **Corpus Redundancy:** Near-duplicate documents (draft iterations, repeated legal disclaimers, email chains) bloat the vector index, diluting retrieval relevance.

- **Compliance Violations:** Hardcoded credentials, internal API keys, customer names, and PII slip into training sets and embedding databases.

- **Context Fragmentation:** Arbitrary fixed-length character chunking severs tables and sentences midway, divorcing specific statements from their qualifying sections.

## The Four-Stage Production Pipeline

### Stage 1: Layout-Aware Structural Normalization

Raw artifacts are first parsed through a vision-assisted document layout engine. Rather than treating files as raw text streams, documents are segmented into structural elements: document titles, H1-H4 headings, paragraph blocks, tables, and caption items. Running headers, page numbers, and boilerplate legal disclaimers are programmatically stripped.

### Stage 2: Scalable Deduplication with MinHash LSH

To identify both exact and near-duplicate documents across millions of records without an O(N²) comparison overhead, the pipeline implements MinHash with Locality-Sensitive Hashing (LSH):

- **Shingling:** Text is tokenized into 5-word character n-grams.

- **MinHash Fingerprinting:** 128 independent hash functions compute dense signature vectors.

- **LSH Banding:** Signatures are partitioned into bands to quickly identify pairs exceeding a 0.85 Jaccard similarity threshold.

This stage removed over 573,000 near-duplicate artifacts (23.9% corpus reduction), drastically reducing downstream embedding computation and vector database footprint.

### Stage 3: Two-Tier Automated PII Sanitization

Data privacy compliance requires zero-tolerance leakage mitigation while preserving syntactic readability. We implemented a two-tier hybrid sanitization engine:

- **Tier 1 (Deterministic Regex):** High-throughput detection of high-entropy strings, API tokens, IP addresses, credit cards, and social security identifiers.

- **Tier 2 (Contextual NER via Microsoft Presidio):** SpaCy-backed transformer entities detecting contextual person names, phone numbers, and enterprise customer identifiers, masking them into normalized canonical tokens (e.g., [PERSON_1], [CLIENT_ORG]).

### Stage 4: Hierarchical Layout-Aware Chunking & Breadcrumb Prepending

Standard fixed-size chunking (e.g., 512 tokens with 50-token overlap) causes catastrophic context loss. We replaced it with semantic boundary chunking combined with *Breadcrumb Prepending*:

Each chunk is prefixed with its document-level taxonomy and parent heading path before embedding generation:

[Document: 2025 Architecture Strategy > Section: Disaster Recovery > Subsection: Failover SLAs]
In the event of primary region disruption, cross-region failover initiates within 120 seconds...

By guaranteeing that every embedding contains its semantic hierarchy, dense retrieval vectors maintain precise contextual alignment even when chunk text is concise.

## Empirical Results & Evaluation Benchmarks

The updated pipeline was evaluated against a held-out benchmark suite of 1,200 enterprise domain queries across our RAG evaluation testbed:

Metric
Naive Scraped Baseline
Production 4-Stage Pipeline
Delta

**Total Artifacts Processed**
2,400,000
1,827,000 (after dedup)
-23.9% Index Size

**Top-5 Context Recall**
61.2%
84.7%
+23.5%

**RAG Faithfulness / Hallucination-Free**
72.4%
91.8%
+19.4%

**PII Leakage Incidents**
142 detected
0 detected
100% Mitigation

**Vector DB Ingestion Latency**
38.4 hours
14.1 hours
-63.3% Cost & Time

## Key Takeaways & Engineering Post-Mortem

- **Data quality eclipses model parameter size:** Enhancing the dataset preparation pipeline delivered larger gains in context recall and factual faithfulness than upgrading model size from 8B to 70B parameters on raw data.

- **Structural breadcrumbs solve needle-in-haystack retrieval:** Prepending document hierarchy anchors dense embeddings to high-level entities without exceeding context budgets.

- **Deduplication is a prerequisite for RAG:** Near-duplicate chunks dilute top-k similarity search, starving generative LLMs of diverse, informative context chunks.

- **Enforce deterministic lineage:** Storing cryptographic hash manifests for every transformation stage ensures pipeline runs are audit-ready, reproducible, and verifiable.

---
### About the Author
**Acadify Engineering Team**
Acadify Engineering Team is the technical team behind Acadify Solution’s AI, software engineering, cloud, automation, and product development work. We publish practical, research-informed insights based on our engineering experience across AI systems, LLM applications, software development, cloud infrastructure, automation, AI testing and evaluation, and digital product engineering. Our content is designed to help founders, engineering teams, technology leaders, and businesses understand complex technical topics and make informed decisions about building, deploying, and improving software and AI systems.
