RAG and search

Retrieval over your own documents or data, answer engines, and natural-language-to-SQL starters.

azure-search-openai-demo leads with 56, ahead of rag-postgres-openai-python (51) and llm-app (48). 12 projects ranked by score.

The ranking

RAG and search: full ranking
RankProjectAdoptionFreshnessMaintenanceEasy to runAgent-readyScore
1azure-search-openai-demoAzure RAG chat reference on AI Search and Azure OpenAI7.8k stars, MIT, last commit Oct 2026731009203056 out of 100
2rag-postgres-openai-pythonRAG over Postgres table rows with hybrid search and SQL filters505 stars, MIT, last commit Oct 20263510049333051 out of 100
3llm-appPathway RAG pipeline templates that re-index live data sources Live demo ↗ (opens in a new tab)58.8k stars, MIT, last commit Jul 20268997330048 out of 100
#4chat-langchainLangChain docs assistant as a Managed Deep Agent with Next.js UI6.5k stars, MIT, last commit Oct 202667100450045 out of 100
#5chat-with-your-data-solution-acceleratorAzure RAG chat app that answers from your documents with citations1.2k stars, MIT, last commit Oct 202644100680044 out of 100
#6llm-answer-enginePerplexity-style Next.js answer engine over Brave search results5k stars, MIT, last commit Apr 20266173033041 out of 100
#7azure-search-openai-javascriptTypeScript RAG on Azure AI Search with separate indexer and search services322 stars, MIT, last commit Sep 20261810010333041 out of 100
#8nextjs-openai-doc-searchBuild-time embeddings of your MDX docs into Supabase pgvector1.7k stars, Apache-2.0, last commit May 20265278033040 out of 100
#9SupabaseAuthWithSSRClaude chat on Next.js 16 with Supabase auth, pgvector RAG, cost dashboards Live demo ↗ (opens in a new tab)397 stars, MIT, last commit Oct 2026291002607039 out of 100
#10natural-language-postgresNext.js text-to-SQL over Postgres with auto-picked charts Live demo ↗ (opens in a new tab)326 stars, Apache-2.0, last commit Apr 20262372033032 out of 100
#11ai-starter-kitSambaNova Python kits for document RAG, search assistant, function calling250 stars, Apache-2.0, last commit Oct 20261185500030 out of 100
#12openai-support-agent-demoSupport console where the model drafts and a human approves203 stars, MIT, last commit Dec 20255240007 out of 100

Momentum, Verified build, Docs and Privacy are not measured yet; their weight goes to the signals shown. A dash means the signal is not scored for that kind of project. Hover a number for its rating in words.

Reviews

156 out of 100

azure-search-openai-demo

Azure RAG chat reference on AI Search and Azure OpenAI

7.8k stars, MIT, last commit Oct 2026

The canonical Azure RAG sample: a Python (Quart) backend and React frontend answering multi-turn questions over your documents with citations and a visible thought process, using Azure AI Search for retrieval and Azure OpenAI for generation. azd up provisions Container Apps, AI Search, Document Intelligence and Blob storage, with optional Cosmos DB chat history, Entra login with document ACLs, multimodal and speech. For teams already on Azure.

Strengths

  • Optional Entra login with per-document access control and Cosmos DB chat history
  • Evaluation, safety evaluation, monitoring and productionizing guides in docs/
  • Multimodal, speech and agentic retrieval are switchable features
  • Commits within the last week; tests included

Weaknesses

  • Cannot run locally until azd up has provisioned Azure resources
  • Provisions paid services by default (AI Search, Document Intelligence); run azd down
  • Azure OpenAI only; no other provider path
  • README itself says not production-ready without extra security work
  • Python, azure-openai
  • Needs azure-subscription, azd, azure-ai-search, azure-openai, azure-document-intelligence, azure-blob-storage
251 out of 100

rag-postgres-openai-python

RAG over Postgres table rows with hybrid search and SQL filters

505 stars, MIT, last commit Oct 2026

A FastAPI backend and React frontend that answer chat questions about rows in a PostgreSQL table. Retrieval is hybrid (pgvector similarity plus full-text search fused with RRF), an OpenAI function call turns phrases like cheaper than 30 dollars into WHERE clauses, and it runs against Azure OpenAI, OpenAI.com or Ollama; azd deploys it to Container Apps with managed identity. For teams whose knowledge is structured rows, not documents.

Strengths

  • Hybrid vector plus full-text search with RRF is implemented in SQL, not a vendor service
  • Provider switch by env var: Azure OpenAI, OpenAI.com or Ollama
  • Evaluation, safety evaluation and load-testing docs included
  • Tests included; dev container and Codespaces configs

Weaknesses

  • Deploy path is Azure-only (azd, Container Apps, Flexible Server)
  • Local run expects Postgres 14+ with pgvector installed yourself
  • Sample schema is one products table; multi-table questions need new code
  • No auth in the app itself
  • Python, azure-openai, openai, ollama
  • Needs postgres-pgvector, azure-openai-or-openai-or-ollama, azd
  • GitHub template
  • env example file
348 out of 100

llm-app

Pathway RAG pipeline templates that re-index live data sources

Live demo ↗ (opens in a new tab)58.8k stars, MIT, last commit Jul 2026

Eight Dockerized Python pipelines on the Pathway framework: question-answering RAG, a live document indexer, multimodal RAG with GPT-4o, unstructured-to-SQL, adaptive RAG, a private Mistral plus Ollama variant, slide search and video RAG. Each watches a source (file system, Google Drive, SharePoint, S3, Kafka, Postgres), keeps an in-memory vector and full-text index current and serves an HTTP API. For teams whose documents change constantly.

Strengths

  • No separate vector DB, cache or API framework; indexing is in-process (usearch, Tantivy)
  • Connectors for file system, Google Drive, SharePoint, S3, Kafka and Postgres with live sync
  • Private variant runs fully local with Mistral and Ollama
  • Docker images and tests included

Weaknesses

  • Pathway is the real dependency; its Rust engine is opaque to most Python teams
  • Pipelines are backends; the UI is an optional Streamlit demo
  • Root README has no setup; each template README is required reading
  • Index lives in memory; sizing for millions of pages is on you
  • Jupyter Notebook, pathway, openai, mistral, ollama
  • Needs docker, openai-api-key, data-source-credentials
#445 out of 100

chat-langchain

LangChain docs assistant as a Managed Deep Agent with Next.js UI

6.5k stars, MIT, last commit Oct 2026

A documentation assistant for LangChain, LangGraph and LangSmith: a Python agent built with LangChain middleware (guardrails, ingress guards, retry) and deployed through Managed Deep Agents, which owns identity, ingress and the checkpointer. Tools search the docs through a managed MCP connector, a Pylon support knowledge base and a URL validator, and a Next.js chat UI sits in frontend/. For teams wanting a reference for a guarded docs assistant.

Strengths

  • Guardrails and link validation are implemented as reusable middleware
  • Supabase token plus guest identity handled in identity.py
  • Frontend proxies LangSmith feedback so the API key never reaches the browser
  • Tests included

Weaknesses

  • Tied to Managed Deep Agents (mda CLI) for identity, ingress and state
  • Needs a Pylon account and knowledge base ID to run as written
  • Docs retrieval depends on a managed MCP connector, not your own index
  • Product-specific: you replace the LangChain docs with your own corpus
  • TypeScript, anthropic, langchain, langgraph
  • Needs anthropic-api-key, pylon-api-key, managed-deep-agents, supabase
  • env example file
#544 out of 100

chat-with-your-data-solution-accelerator

Azure RAG chat app that answers from your documents with citations

1.2k stars, MIT, last commit Oct 2026

Deploys a React frontend, a FastAPI backend and an Azure Functions ingestion worker to Azure Container Apps with `azd up`. Uploaded files and web pages are parsed, chunked and embedded, then answered with streamed responses and inline citations. Retrieval and chat history use either Azure AI Search with Cosmos DB or PostgreSQL with pgvector, chosen at deploy time.

Strengths

  • Choice of Azure AI Search + Cosmos DB or PostgreSQL + pgvector at deploy time
  • Managed identity and RBAC for all calls; no Key Vault or app secrets
  • Admin UI for ingesting documents and editing prompts without code changes
  • Two selectable orchestrators: Agent Framework or LangGraph

Weaknesses

  • Azure-only; requires Foundry, Document Intelligence, Storage, and Container Apps
  • Needs Contributor and RBAC rights on the subscription, plus model quota
  • README calls it a starting point, not production-ready
  • No Docker or compose files detected in the repo; local setup is in docs
  • Python, Azure AI Foundry chat and embedding models
  • Needs Azure AI Foundry, Azure AI Search, Azure Cosmos DB, Azure Database for PostgreSQL, Azure Document Intelligence, Azure Storage, Azure Container Apps, Azure Functions, Azure Content Safety, Azure AI Speech
  • Docker
  • env example file
#641 out of 100

llm-answer-engine

Perplexity-style Next.js answer engine over Brave search results

5k stars, MIT, last commit Apr 2026

A Next.js app that takes a question, pulls results from Brave Search and Serper, scrapes the top pages with Cheerio, chunks and embeds them with OpenAI embeddings, and streams an answer from Groq (Mixtral by default) with sources, images and follow-ups. Optional Ollama, Upstash rate limiting, a semantic cache and a Portkey gateway are toggles in app/config.tsx; there is no auth or persistence. For developers learning the search-scrape-answer loop.

Strengths

  • Full pipeline readable in one config file: search, scrape, chunk, embed, answer
  • docker compose and a standalone Express API variant included
  • Optional rate limiting and semantic cache via Upstash

Weaknesses

  • Four API keys to start (OpenAI, Groq, Brave, Serper)
  • No auth, no chat history, no tests
  • Pinned to Next.js 14.1 and dated defaults (mixtral-8x7b-32768)
  • Ollama mode skips follow-up questions; vectors are in-memory only
  • TypeScript, groq, openai, ollama, portkey
  • Needs openai-api-key, groq-api-key, brave-search-api-key, serper-api-key
  • Docker
  • env example file
#741 out of 100

azure-search-openai-javascript

TypeScript RAG on Azure AI Search with separate indexer and search services

322 stars, MIT, last commit Sep 2026

The Node.js counterpart of the Azure RAG sample: a search API, an indexer service and a web app that answer chat and Q&A questions over your documents with citations, using Azure AI Search and Azure OpenAI through LangChain.js. azd up provisions Container Apps for the backend and a Static Web App for the frontend, and the search API speaks the AI chat HTTP protocol so the Python backend can replace it. For TypeScript teams on Azure.

Strengths

  • Indexer, search API and web app are separate services with their own deploys
  • Search API follows the AI chat HTTP protocol; the backend is swappable
  • Tests included; Codespaces and dev container configs

Weaknesses

  • Cannot run locally until azd up has provisioned Azure resources
  • No authentication shipped; Entra setup is a linked tutorial
  • Azure OpenAI and Azure AI Search only
  • Less active than the Python sample (seed stars 322 vs 7776)
  • TypeScript, azure-openai, langchain
  • Needs azure-subscription, azd, azure-ai-search, azure-openai, azure-blob-storage
  • GitHub template
#939 out of 100

SupabaseAuthWithSSR

Claude chat on Next.js 16 with Supabase auth, pgvector RAG, cost dashboards

Live demo ↗ (opens in a new tab)397 stars, MIT, last commit Oct 2026

A Next.js 16 app on AI SDK v7 and Claude with complete Supabase SSR auth (signup, magic links, password reset, RLS on every table) and eight tools: PDF RAG (Mistral OCR, Voyage embeddings, hybrid RRF search in pgvector), Exa web search, versioned artifacts, memory, conversation search, sandboxed visualizations, PDF export and image generation. Per-step token usage feeds user and admin cost dashboards. For teams shipping a paid Claude assistant on Supabase.

Strengths

  • Whole schema, RLS, triggers and search functions in one idempotent setup.sql
  • Per-step token and cache usage stored on messages; user and admin cost dashboards
  • Two-tier Anthropic prompt caching with a hit-rate readout
  • Live instance at supa-chat.dev; committed within the last day

Weaknesses

  • Anthropic-only chat; OCR, embeddings and search add Mistral, Voyage and Exa keys
  • No tests listed
  • Image generation needs your own GPU server (RTX 5090 32 GB recommended)
  • One maintainer; large surface area to understand before customizing
  • TypeScript, anthropic, ai-sdk, mistral, voyage
  • Needs supabase, anthropic-api-key, mistral-api-key, voyage-api-key, exa-api-key
  • env example file
  • sign-in: Supabase Auth
#1032 out of 100

natural-language-postgres

Next.js text-to-SQL over Postgres with auto-picked charts

Live demo ↗ (opens in a new tab)326 stars, Apache-2.0, last commit Apr 2026

A Next.js app where the AI SDK and GPT-4o turn a plain-English question into SQL, run it against Postgres, show the rows, pick a chart type and render it with Recharts, and explain the query on request. It ships with a seed script for a unicorn-companies CSV you download yourself; there is no auth and no history. For developers who want a text-to-SQL and charting pattern to copy.

Strengths

  • Shows the full loop: generate SQL, execute, explain, chart config, render
  • Deployed demo on Vercel
  • Two secrets to run: OPENAI_API_KEY and a Postgres URL

Weaknesses

  • Single hardcoded dataset; the schema prompt must be rewritten for your tables
  • OpenAI GPT-4o only
  • No auth, tests or history
  • Dataset CSV must be fetched manually from CB Insights
  • TypeScript, openai, ai-sdk
  • Needs postgres, openai-api-key
  • env example file
#1130 out of 100

ai-starter-kit

SambaNova Python kits for document RAG, search assistant, function calling

250 stars, Apache-2.0, last commit Oct 2026

Nine Python kits, each with its own README: document text extraction, enterprise and multimodal knowledge retrieval with Streamlit demos, a RAG evaluation kit, a web search assistant, a financial assistant using function calling and scraping, a function-calling module, benchmarking and chat templates. Everything calls SambaNova models through SAMBANOVA_API_KEY. For teams on SambaCloud or SambaStack who want working retrieval code.

Strengths

  • Knowledge retriever and search assistant kits include runnable Streamlit demos
  • Makefile base environment installs Python, Poetry, Tesseract and Poppler; Docker option
  • RAG evaluation kit included

Weaknesses

  • SambaNova endpoints only; swapping providers means editing each kit
  • README states the code is as-is and not production-ready
  • Mixed notebooks and apps; no single app to fork
  • Heavy setup: pyenv, Poetry, a parsing service and OCR system packages
  • Jupyter Notebook, sambanova, langchain
  • Needs sambanova-api-key, tesseract, poppler
  • Docker
#127 out of 100

openai-support-agent-demo

Support console where the model drafts and a human approves

203 stars, MIT, last commit Dec 2025

A Next.js demo on the OpenAI Responses API with two chat views, one for the customer and one for the human agent. The model drafts replies from a file-search knowledge base, proposes tool calls like cancel_order for the agent to confirm and auto-runs non-sensitive ones like get_order_history; a /init_vs route creates the vector store and functions are placeholders. For teams prototyping agent-assist for support staff.

Strengths

  • Human-in-the-loop pattern is concrete: suggested reply, suggested action, auto-run tiers
  • Knowledge base, prompts, tools and demo data each live in one config file
  • File search vector store bootstrapped from a route

Weaknesses

  • README says not production-ready: no auth, no guardrails
  • Tool functions are stubs that change nothing
  • OpenAI Responses API only
  • Last commit 2025-12
  • TypeScript, openai
  • Needs openai-api-key
  • env example file

Written from each project's README and checked facts. Spot something wrong? Report it on GitHub (opens in a new tab).