Chunk smart. Embed right. Retrieve what the agent actually needs.
AI Agentic Practice
Data Chunking & Embeddings
Poor chunking and embeddings are the #1 reason enterprise RAG fails. Gensten designs chunking strategies (fixed, recursive, semantic, layout-aware), embedding model selection, dimensionality and refresh policies, ACL metadata, and index lifecycle - so retrieval stays precise as your OneLake and document corpus scale.
Model choice, batching, versioning, quality checks.
5
Index & Maintain
Upserts, deletes, drift detection, A/B of embedding versions.
Chunking is architecture, not a parameter
Token size alone is not a strategy. Legal contracts, support tickets, and codebases need different splits. We prototype chunkers against golden Q&A sets and measure retrieval lift before locking production settings.
Parent documents + child chunks for context expansion
Table and figure-aware splitting
Language-specific tokenizers
Deduplication and near-duplicate collapse
Embedding strategy
We compare embedding families for multilingual, code, and domain jargon; track version pins; and plan migrations so agents don’t silently degrade when you change models.
It depends on document structure and query style. We run empirical tests (often 256–1024 tokens with overlap, plus semantic splits) against your evaluation set rather than guessing a global default.
What is data chunking in RAG?
Chunking splits documents into retrieval units. Good chunks preserve meaning, keep metadata/ACLs, and fit the model context so the LLM gets the right evidence - not random fragments.
What are embeddings?
Embeddings are numeric vectors that represent meaning. Similar text maps to nearby vectors, enabling semantic search over policies, tickets, code, and knowledge bases.
Fixed-size vs. semantic chunking - which is better?
Fixed-size is simple; semantic/layout-aware chunking usually wins for PDFs, HTML, and structured docs. Many production systems combine hierarchical parent-child chunks.
How much overlap should chunks have?
Moderate overlap (e.g., 10–20%) reduces boundary misses. Too much overlap wastes storage and can dilute ranking. We tune overlap with retrieval metrics, not rules of thumb alone.
How do you choose an embedding model?
We benchmark domain Q&A recall, multilingual needs, code vs. prose, dimension/cost tradeoffs, and migration risk - then pin versions so indexes stay consistent.
What metadata should every chunk store?
Source ID/URL, title, section, timestamps, language, sensitivity label, and ACL principals. Metadata powers filtering, citations, and compliance.
How do you handle tables and images?
Tables are extracted to structured text or row-aware chunks; figures may use captions/OCR. Layout-aware parsing is critical for financial and technical PDFs.
When should we re-embed the corpus?
When changing embedding models, major content restructure, or measured retrieval drift. We maintain versioned indexes and A/B before cutting over.
Can chunking pipelines run on OneLake?
Yes. Source docs and derived chunk/embedding datasets can live in OneLake with lineage, feeding Azure AI Search or other vector indexes for Foundry agents.