Hero background

DATA CHUNKING & EMBEDDINGS

Chunk smart. Embed right. Retrieve what the agent actually needs.

AI Agentic Practice

Data Chunking & Embeddings

Poor chunking and embeddings are the #1 reason enterprise RAG fails. Gensten designs chunking strategies (fixed, recursive, semantic, layout-aware), embedding model selection, dimensionality and refresh policies, ACL metadata, and index lifecycle - so retrieval stays precise as your OneLake and document corpus scale.

  • Document-type aware chunking (PDF, HTML, code, tables)
  • Semantic and hierarchical chunk strategies
  • Embedding model benchmarks for your domain
  • Metadata, ACLs, and lineage for every chunk
  • Re-embed / reindex playbooks as models improve
Data Chunking & Embeddings

Architecture Layers

How Chunking & Embeddings fits into a production agentic stack.

1

Parse & Normalize

OCR, HTML cleaning, table extraction, language detection.

2

Chunk

Size, overlap, semantic boundaries, parent-child hierarchies.

3

Enrich

Titles, summaries, entities, sensitivity labels, source URLs.

4

Embed

Model choice, batching, versioning, quality checks.

5

Index & Maintain

Upserts, deletes, drift detection, A/B of embedding versions.

Chunking is architecture, not a parameter

Token size alone is not a strategy. Legal contracts, support tickets, and codebases need different splits. We prototype chunkers against golden Q&A sets and measure retrieval lift before locking production settings.

  • Parent documents + child chunks for context expansion
  • Table and figure-aware splitting
  • Language-specific tokenizers
  • Deduplication and near-duplicate collapse

Embedding strategy

We compare embedding families for multilingual, code, and domain jargon; track version pins; and plan migrations so agents don’t silently degrade when you change models.

Use Cases

  • Large SharePoint / Confluence migrations into RAG
  • Multi-tenant knowledge indexes
  • Codebase RAG for developer agents
  • OneLake document lakes with vector indexes

Technologies & Platforms

OpenAI / Azure embeddingsSentence TransformersUnstructured.io / custom parsersAzure AI SearchpgvectorOneLake shortcuts

Frequently Asked Questions

What chunk size should we use?

It depends on document structure and query style. We run empirical tests (often 256–1024 tokens with overlap, plus semantic splits) against your evaluation set rather than guessing a global default.

What is data chunking in RAG?

Chunking splits documents into retrieval units. Good chunks preserve meaning, keep metadata/ACLs, and fit the model context so the LLM gets the right evidence - not random fragments.

What are embeddings?

Embeddings are numeric vectors that represent meaning. Similar text maps to nearby vectors, enabling semantic search over policies, tickets, code, and knowledge bases.

Fixed-size vs. semantic chunking - which is better?

Fixed-size is simple; semantic/layout-aware chunking usually wins for PDFs, HTML, and structured docs. Many production systems combine hierarchical parent-child chunks.

How much overlap should chunks have?

Moderate overlap (e.g., 10–20%) reduces boundary misses. Too much overlap wastes storage and can dilute ranking. We tune overlap with retrieval metrics, not rules of thumb alone.

How do you choose an embedding model?

We benchmark domain Q&A recall, multilingual needs, code vs. prose, dimension/cost tradeoffs, and migration risk - then pin versions so indexes stay consistent.

What metadata should every chunk store?

Source ID/URL, title, section, timestamps, language, sensitivity label, and ACL principals. Metadata powers filtering, citations, and compliance.

How do you handle tables and images?

Tables are extracted to structured text or row-aware chunks; figures may use captions/OCR. Layout-aware parsing is critical for financial and technical PDFs.

When should we re-embed the corpus?

When changing embedding models, major content restructure, or measured retrieval drift. We maintain versioned indexes and A/B before cutting over.

Can chunking pipelines run on OneLake?

Yes. Source docs and derived chunk/embedding datasets can live in OneLake with lineage, feeding Azure AI Search or other vector indexes for Foundry agents.

Ready to build production AI agents?

Talk to Gensten about ADLC, RAG, LLM building, Foundry IQ, Fabric IQ, and OneLake - scoped to your KPIs and compliance needs.