Atlas — Agentic RAG Platform
Multi-tenant LLM platform with namespace-isolated corpora, a 5-stage agentic retrieval pipeline, and full cost/telemetry instrumentation. Live on GCP.
Atlas is a multi-tenant agentic RAG platform — the kind of retrieval backend a product team would build an AI feature on top of. Built in async Python (~6,100 LOC) and deployed on Google Cloud Run in asia-south1, it is live at atlas.hulage.in.
The problem
Most RAG demos are single-user notebooks. Turning one into something a team can actually depend on means solving the boring, load-bearing parts: keeping tenants isolated, authenticating callers, capping spend, and proving that a change to the pipeline actually improved retrieval instead of quietly regressing it.
How it works
- Multi-tenancy & security — namespace-isolated corpora so one tenant can never read another’s data, SHA-hashed API-key auth, per-key rate limiting, token-level cost attribution, and Prometheus telemetry on every request.
- Agentic retrieval — a 5-stage pipeline that fuses dense vector search over Qdrant with BM25 keyword search via Reciprocal Rank Fusion, then a cross-encoder reranks the merged candidates before generation.
- Measured, not guessed — an evaluation harness runs a labeled question set against the pipeline so retrieval changes are A/B-tested on precision and recall rather than eyeballed.
- Idempotent ingestion — content fingerprinting skips unchanged chunks, so re-indexing a corpus never duplicates or corrupts the vector store.
Stack
Gemini via an OpenAI-compatible base URL, Qdrant Cloud for vectors, Firestore for API keys and spend tracking, FastAPI + Docker on Cloud Run, fronted through a Vercel rewrite. The whole thing runs inside a hard daily cost cap.