June 15, 2026

SourceFetch: Grounded Q&A for Billing and Coding Policy

A retrieval augmented assistant that answers billing and coding questions using only the source manuals. It cites the relevant passages and clearly says when the answer is not supported by the available sources.

  • python
  • fastapi
  • rag
  • chromadb
  • llm
  • healthcare

Launch Live Demo

Billers and coders spend their day working through dense policy PDFs. The corpus includes the NCCI Policy Manual, the Claims Processing Manual, and modifier guidance. They search for the one paragraph that settles a question. SourceFetch turns that corpus into a question answering service. It retrieves relevant passages from a vector store and uses a model to answer strictly from those passages. Each passage is cited inline. SourceFetch is the companion to ScrubCheck. ScrubCheck flags that a claim will deny. SourceFetch explains why based on the manual.

The code is available on GitHub. The design is summarized below.

System Design

SourceFetch uses retrieval augmented generation (RAG). The governing rule is that the model may speak only from the passages it retrieves. A question is embedded, the most relevant passages are retrieved from a ChromaDB vector store, and a model answers strictly from those passages. Each claim is cited with a [n] that links to the exact source. I built the system around retrieval rather than fine tuning the model on the policy because the rulebook changes every quarter. An answer that traces back to a specific passage is more trustworthy than one baked opaquely into a model’s weights. The tool runs offline without an API key and returns the raw retrieved passages. When configured, it can use Claude to generate the final prose and Hugging Face models for semantic embeddings.

Enforcing Abstention

The hardest part was getting the model to refuse. When asked about a modifier that was not covered in the loaded manuals, an early version would answer from its own background knowledge or stretch a nearby passage to fit. That is exactly the kind of failure a billing tool cannot have. I tightened the grounding requirements so the model abstains when the retrieved sources do not support an answer and clearly states what the corpus does cover. This turned the tool’s most serious vulnerability into one of its strongest features.

SourceFetch declining to answer a question about a code not covered by its sources.

Evaluation

Most RAG demos ship without meaningful evaluation. This one includes an evaluation harness built around a small gold set. The harness measures retrieval quality using hit@k and MRR, along with a grounding check that uses an LLM as a judge. The question, “Is it any good?” now has a numeric answer rather than relying on impression. I led the evaluation workstream for my graduate capstone, and that experience shaped this approach. Evaluation is what separates a working RAG demo from a clear account of how well it performs and where it fails.

Deployment and Hosting

The retrieval layer is inexpensive and safe to expose. SourceFetch is containerized and runs as a public demo on Google Cloud Run in retrieval-only mode. The answer generation layer uses Claude and makes paid API calls, so it remains gated behind an API key. As with the rest of my work, the source code is available on GitHub, and the demo above shows it running.

Receipts

← All projects