hRAG logo

hRAG

Retrieval with receipts

usage based Any browser Kubernetes AI Search

hRAG is a ai search tool. It's best for Developers and Organizations seeking self-hosted RAG solutions. Pricing is usage based.

Pricing

usage based

Audience

Developers

Platforms

Community

0%

About hRAG

hRAG is a self-hosted hybrid Retrieval Augmented Generation (RAG) platform designed to run on a low-cost Kubernetes cluster, offering BM25, vector search, and a reranker for accurate information retrieval.

hRAG provides a self-hosted hybrid RAG solution that integrates BM25 lexical search and vector embeddings within a single Postgres database. This architecture eliminates the need for separate search clusters, simplifying operations and reducing infrastructure costs. The platform includes a cross-encoder reranker to improve retrieval accuracy, which is offered as an optional feature due to its latency impact.

The system emphasizes transparency and verifiability, with every claim backed by public benchmark scores and reproducible results. It features grounded answers, refusing to hallucinate when information is not found, and provides streaming citations that link directly to source chunks. Tenant isolation is enforced using Postgres row-level security, ensuring data privacy for multiple users or sandboxes.

hRAG is designed for cost-efficiency, running on a five-node Kubernetes cluster on Hetzner for approximately €116 per month. It utilizes small, budget-friendly models for embedding and answering, benchmarked against more expensive cloud alternatives. The entire codebase is MIT licensed, and deployment articles are provided to guide users through self-hosting.

The platform is officially scored on EnterpriseRAG-Bench, achieving a #9 ranking with strong performance in correctness, completeness, and document recall, outperforming Vertex AI Search and NVIDIA in some metrics. Users can try a 512K-document playground without login, or sign in with Google/GitHub for a private sandbox with limited document capacity and token budget.

Key Features

Self-hosted hybrid RAG
BM25 lexical search within Postgres
Vector embeddings within Postgres
Cross-encoder reranker (optional)
Grounded answers with refusal for empty retrieval
Streaming citations linked to source chunks
Tenant isolation via Postgres row-level security
Public benchmark scores and reproducible results
Cost-efficient infrastructure (€116/month cluster)
MIT licensed code
512K-document playground
Private sandbox for own documents (Google/GitHub login)
Small, budget-friendly AI models

Pricing

usage based

The entire five-node Kubernetes cluster costs €116 per month on Hetzner. Answers cost tenths of a cent. The benchmark that proved the system cost about $60 once. A free playground is available without login. A private sandbox is available upon Google or GitHub sign-in, offering 10 documents (20 pages each) and a daily token budget.

Who is it for?

Best for

  • Organizations prioritizing data privacy and self-hosting for RAG
  • Developers and teams looking for a transparent, benchmarked RAG solution
  • Cost-conscious users who want to run RAG on affordable infrastructure
  • Research and development into RAG systems

Not ideal for

  • Users seeking a fully managed, zero-setup RAG service
  • Organizations unwilling to manage their own Kubernetes cluster
  • Users requiring extensive, out-of-the-box integrations with various SaaS platforms

Integrations

Postgres Google GitHub

Community Discussion

Sign in to contribute

No discussions yet. Be the first to share your experience!

Frequently asked questions