Accurate, citeable answers from your documents, lakes, and systems of record.
AI Agentic Practice
Enterprise RAG Systems
Gensten builds enterprise RAG systems that ground LLMs on your private data. We design ingestion pipelines, data chunking strategies, embedding models, hybrid search, rerankers, citation UX, and evaluation - on Azure AI Search, Fabric/OneLake, or open vector databases - so agents and copilots answer with evidence, not hallucinations.
Agents fail when retrieval is weak. We treat retrieval quality as a product: measure recall@k, citation coverage, and groundedness, then feed results back into chunking and embedding choices.
Per-domain indexes with ACL-aware filtering
Query expansion and tool-assisted retrieval
Freshness SLAs for knowledge corpora
Separation of public vs. confidential collections
Microsoft-aligned knowledge paths
For Microsoft-centric estates we wire RAG to OneLake and Fabric IQ so analytics and unstructured knowledge share one governed lakehouse - with Foundry IQ agents retrieving through certified tools.
Use Cases
Policy & procedure copilots
Customer support knowledge agents
Engineering runbook assistants
Legal / contract Q&A with citations
Technologies & Platforms
Azure AI SearchOneLakeFabricOpenAI embeddingsLangChain / LlamaIndexPinecone / Weaviate / pgvector
Frequently Asked Questions
When should we use RAG vs. fine-tuning?
Use RAG when facts change often or live in enterprise documents. Fine-tune for style, domain language, or task format. Most enterprises need both: RAG for knowledge + light adaptation for behavior.
What is Retrieval-Augmented Generation (RAG)?
RAG retrieves relevant enterprise content at query time and feeds it to an LLM so answers are grounded in your data with citations - reducing hallucinations versus prompting alone.
Why do enterprise RAG systems fail?
Common causes: poor chunking, weak embeddings, missing ACLs, stale indexes, no reranking, and no evaluation. We treat retrieval quality as a product metric, not a one-time setup.
Do you support hybrid search?
Yes. We combine keyword and vector search with query rewrite and reranking so exact terms (IDs, codes, names) and semantic matches both surface correctly.
How do citations work in RAG?
Retrieved chunks carry source metadata (document, URL, page, ACL). The generation layer is prompted to cite sources, and we evaluate citation coverage in the test harness.
Can RAG respect document permissions?
Yes. We propagate ACLs into the index and filter retrieval by user/group identity so agents only see content the requester is allowed to access.
Which vector stores do you use?
Azure AI Search for Microsoft estates, plus Pinecone, Weaviate, or pgvector when preferred. Choice depends on scale, hybrid search needs, and cloud strategy.
How often should the RAG index refresh?
It depends on content volatility. Policies, runbooks, and tickets may need near-real-time or hourly updates; stable corpora can refresh daily with change detection.
Is RAG enough for agentic AI?
RAG is necessary for knowledge grounding but not sufficient alone. Agents also need tools, orchestration, evals, and ADLC operations for production outcomes.
Can RAG run on Microsoft OneLake data?
Yes. We design OneLake document and table layouts, then embed and index for Azure AI Search or other stores so Foundry agents retrieve from governed lake content.