pathwaycom/llm-app

Ready-to-deploy Pathway templates for live-synced RAG and AI enterprise search, with built-in vector, hybrid and full-text indexing, runnable as Docker containers.

  • 58.8k GitHub stars
  • Jupyter Notebook
  • ⚖️ MIT
  • 🎯 Intermediate
pathwaycom/llm-app preview image

What it is

pathwaycom/llm-app is a repo of LLM app templates built on the Pathway Live Data Framework. The templates provide high-accuracy RAG and AI enterprise search over data sources that stay in sync (additions, deletions, updates). They expose an HTTP API, can run as Docker containers, and include built-in in-memory indexing (vector, hybrid, full-text) with no separate infrastructure to set up.

Who it's for

  • Developers who want production-oriented RAG or enterprise search templates over live data
  • Teams syncing documents from file systems, Google Drive, Sharepoint, S3, Kafka, PostgreSQL or real-time data APIs
  • Developers who want to avoid assembling a separate vector database, cache and API framework
  • Teams that need a fully private/local RAG option using Mistral and Ollama

Requirements

Requirements

  • Docker, if running the apps as containers
  • Per-template instructions from each template's README.md in the templates folder
  • A connected data source (e.g. files, Google Drive, Sharepoint)
  • Access to the model or service a given template uses (e.g. GPT models, GPT-4o, TwelveLabs for the video template; Mistral and Ollama for the private template)

Setup

  1. Choose a template and follow its README

    Each template in the templates/ folder contains a README.md with run instructions. You can test a template on your own machine and then deploy to cloud (GCP, AWS, Azure, Render, ...) or on-premises. The main README gives no single install command.

Examples

Question-Answering RAG App

Prompt
prompt
templates/question_answering_rag/

Expected output: Basic end-to-end RAG app that answers questions about documents (PDF, DOCX, ...) on a live connected data source using the GPT model of your choice.

Live Document Indexing as a retriever backend

Prompt
prompt
templates/document_indexing/

Expected output: Real-time indexing pipeline acting as a vector store service; it works with any frontend or as a retriever backend for Langchain or Llamaindex applications.

Unstructured-to-SQL pipeline

Prompt
prompt
templates/unstructured_to_sql_on_the_fly/

Expected output: Structures financial report PDFs into a PostgreSQL table and answers natural language questions by translating them into SQL with an LLM.

Private RAG with Mistral and Ollama

Prompt
prompt
templates/private_rag/

Expected output: A fully private, local version of the question_answering_rag pipeline.

Pros & cons

Pros

  • Pro:Built-in indexing (vector via usearch, hybrid full-text via Tantivy) means no separate vector database, cache or API framework to maintain
  • Pro:Data sources stay in sync with additions, deletions and updates across many connectors
  • Pro:Wide template selection, including multimodal, SQL, adaptive RAG, private/local, slides and video options
  • Pro:Docker-friendly, with an HTTP API for connecting any frontend

Cons

  • Con:The main README has no concrete setup commands; you must consult each template's README
  • Con:Depends on the Pathway Live Data Framework, so adopting it means committing to that stack
  • Con:Several templates rely on external model services (e.g. GPT-4o, TwelveLabs), though a private local option exists

Images