Skip to content
Veritas AI

Hybrid retrieval · Grounded generation · Full observability

Build production-grade AI applications with enterprise RAG

Ingest anything, retrieve with hybrid precision, and stream answers your users can verify — every claim cited, every stage observable.

Trusted by teams at

  • Microsoft
  • Google
  • AWS
  • Azure
  • OpenAI
0.00%uptime SLA
<0msp50 retrieval
0M+documents indexed

veritas — rag shell

Powering retrieval for engineering teams worldwide

Features

Everything a production RAG stack needs

Every stage of the pipeline is observable, swappable, and built for scale — no glue code required.

Document Parsing

PDFs, HTML, and office formats parsed with layout awareness — tables, headings, and reading order preserved.

OCR

Scanned pages and embedded images pass through OCR automatically, so nothing in your corpus is invisible.

Smart Chunking

Recursive, semantic, and fixed-window strategies with configurable size and overlap — swappable per collection.

Embeddings

Local sentence-transformers or hosted models behind one interface. Change providers with an env var, not a rewrite.

Hybrid Search

Dense vectors and sparse BM25 run in parallel, fused with Reciprocal Rank Fusion for rankings neither achieves alone.

Reranking

Cross-encoder rerankers re-score the candidate set so the context window only carries what actually matters.

Vector Search

Chroma, Pinecone, or Qdrant behind a single store interface — local-first in development, managed at scale.

Streaming

Tokens stream over SSE the moment generation starts, with sources emitted first so citations render instantly.

Citations

Every claim links back to the exact chunk, page, and score it came from. Answers your users can verify.

Guardrails

Retrieved content is treated as data, never instructions — prompt-injection surface minimized by design.

Evaluation

RAGAS-style faithfulness, relevancy, and context precision scored continuously against your live traffic.

Observability

Per-stage latency, token usage, and cache hit rates for every request — Grafana-grade visibility built in.

Pipeline

Every stage, explorable

From upload to citation — click any stage to see its purpose, best practices, library options, and the tradeoffs that matter.

Models

Bring the models you trust

Embeddings, LLMs, and rerankers are pluggable — compare the options and swap providers with a config change, not a rewrite.

Chunking

See exactly how your documents split

Tune chunk size and overlap and watch the segmentation change — the same math the ingestion pipeline runs.

bank-policy.pdf

6 chunks

Savings accounts require a minimum balance of $300, calculated as the average daily balance across each statement cycle. Students and account holders under the age of 25 are exempt from the minimum balance requirement until their 25th birthday. If an account falls below the minimum for two consecutive statement cycles, a monthly maintenance fee of $5 applies until the balance is restored. Checking accounts carry no minimum balance, but overdraft protection requires a linked savings account in good standing. Wire transfers within the same business day are subject to a $25 domestic fee and a $45 international fee, waived for premium tier members. Certificates of deposit may be withdrawn early with a penalty equal to 90 days of accrued interest on terms of one year or less, and 180 days on longer terms. Premium tier membership requires a combined balance of $25,000 across all linked accounts, reviewed quarterly against the average combined balance.

220 chars
40 chars

blended regions belong to two chunks — overlap keeps answers that span a boundary retrievable.

Observability

See every request, stage by stage

Latency percentiles, per-stage breakdowns, token throughput, and host telemetry — live for every deployment, no extra wiring.

Evaluation

Quality you can measure, not vibe-check

RAGAS-style scoring runs continuously: faithfulness, relevancy, and context quality tracked across every eval run.

0.00

Faithfulness

Claims supported by retrieved context

0.00

Answer relevancy

Answer addresses the question asked

0.00

Groundedness

No content invented beyond sources

0.00

Context precision

Retrieved chunks that were actually useful

0.00

Context recall

Useful chunks that were actually retrieved

RunDatasetFaithfulnessRelevancyPrecisionRecallCost / queryp95
eval-2041banking-qa-v30.960.930.890.91$0.00411.9s
eval-2040banking-qa-v30.940.920.870.90$0.00432.1s
eval-2039support-kb-v10.910.900.840.88$0.00391.8s
eval-2038banking-qa-v20.930.910.860.89$0.00462.3s

Ship AI your users can actually trust

Start free on your own machine. Scale to millions of documents when you're ready. Every answer cited, every stage observable.

Veritas AI · free tier forever · no credit card required