Back to Projects
Live2026

Atlas — Agentic RAG Platform

Multi-tenant LLM platform with namespace-isolated corpora, a 5-stage agentic retrieval pipeline, and full cost/telemetry instrumentation. Live on GCP.

PythonFastAPIQdrantGeminiCloud RunDocker
Live Demo

Atlas is a multi-tenant agentic RAG platform — the kind of retrieval backend a product team would build an AI feature on top of. Built in async Python (~6,100 LOC) and deployed on Google Cloud Run in asia-south1, it is live at atlas.hulage.in.

The problem

Most RAG demos are single-user notebooks. Turning one into something a team can actually depend on means solving the boring, load-bearing parts: keeping tenants isolated, authenticating callers, capping spend, and proving that a change to the pipeline actually improved retrieval instead of quietly regressing it.

How it works

  • Multi-tenancy & security — namespace-isolated corpora so one tenant can never read another’s data, SHA-hashed API-key auth, per-key rate limiting, token-level cost attribution, and Prometheus telemetry on every request.
  • Agentic retrieval — a 5-stage pipeline that fuses dense vector search over Qdrant with BM25 keyword search via Reciprocal Rank Fusion, then a cross-encoder reranks the merged candidates before generation.
  • Measured, not guessed — an evaluation harness runs a labeled question set against the pipeline so retrieval changes are A/B-tested on precision and recall rather than eyeballed.
  • Idempotent ingestion — content fingerprinting skips unchanged chunks, so re-indexing a corpus never duplicates or corrupts the vector store.

Stack

Gemini via an OpenAI-compatible base URL, Qdrant Cloud for vectors, Firestore for API keys and spend tracking, FastAPI + Docker on Cloud Run, fronted through a Vercel rewrite. The whole thing runs inside a hard daily cost cap.