Skip to content

Exam Slayer

An AI-powered exam preparation system that turns course notes and past papers into grounded, marks-aware study packs and solved answer PDFs.

2026 · Solo · Shipped

ReactFastAPIPythonGeminiRAGTesseractDocker

The problem

Exam preparation often means jumping between lecture notes, textbooks, past papers and manually written answers. The challenge is not only finding information, but turning that material into answers that match the question, available course context and marks allocated. Exam Slayer V2 explores an end-to-end workflow that ingests course material and question papers, retrieves relevant context, generates marks-aware answers, evaluates the output, and packages the result into a print-ready study guide.

How it works

  1. Materials ingested

    Course notes, textbooks and past papers are uploaded.

  2. Text extracted

    PDFs are parsed with PyMuPDF/Tesseract, with page rendering available for visual content.

  3. Questions classified

    A deterministic pre-parser and Gemini identify question boundaries and marks.

  4. Context retrieved

    Relevant course material is retrieved using source-aware TF-IDF RAG and supplied to the generation step.

  5. Answers compiled

    Marks-aware answers are evaluated, retried when needed, and rendered into a print-ready PDF.

The pipeline combines deterministic parsing with retrieval and LLM generation rather than relying on a single prompt to process an entire exam paper.

What I built

  • Solo: I built the complete React + FastAPI application and the document-processing pipeline behind it.
  • Built the React dashboard for uploading materials, selecting preparation modes, tracking processing stages and downloading generated study packs.
  • Implemented deterministic question parsing combined with Gemini-based classification to identify question boundaries and marks.
  • Built a source-aware TF-IDF retrieval pipeline that retrieves relevant course material and attaches source citations to generated answers.
  • Added OCR support with Tesseract for scanned exam papers.
  • Added multimodal page processing so diagram-heavy PDF pages can be rendered and supplied to the model as visual context.
  • Implemented marks-aware answer generation for 2-, 5- and 10-mark questions.
  • Added a deterministic quality gate that checks citation presence, answer length and placeholder leakage, with one automatic retry for failed batches.
  • Built the PDF generation pipeline using HTML/CSS templates and WeasyPrint.
  • Containerized and deployed the application through Docker on Hugging Face Spaces.

Decisions

RAG grounding over free answering

Course-specific answers should be based on the material provided by the student rather than relying only on the model's general knowledge. I added local retrieval so relevant sections of uploaded notes are included in the generation context.

Marks-aware solving

A 2-mark answer and a 10-mark answer should not have the same structure or depth. The pipeline uses the question's marks allocation to control answer length and structure.

Deterministic parsing before LLM processing

Instead of sending an entire paper to the model and asking it to infer everything, I added a deterministic parsing stage first. This reduces ambiguity and gives later generation stages cleaner inputs.

Quality gate + retry

Generated answers are checked for citation coverage, marks-aware length and placeholder leakage. Failed batches receive a stricter retry prompt instead of being accepted blindly.

Simple retrieval before adding infrastructure

I deliberately started with local TF-IDF retrieval instead of introducing a vector database. For the current study workflow, the simpler architecture was easier to debug, cheaper to operate and sufficient for the current corpus. A semantic cache or vector store is a future direction if the system needs larger-scale retrieval.

Architecture

Course material / question paper
  → Extraction + OCR
  → Deterministic pre-parser
  → Gemini classification
  → TF-IDF retrieval (RAG)
  → Marks-aware generation
  → Quality gate + retry
  → HTML / PDF renderer
  → Print-ready study pack

The backend is built with FastAPI and Python, while the React frontend manages uploads, progress and generated results. The processing pipeline combines deterministic document parsing, retrieval, Gemini generation and PDF rendering.

Results and evaluation

Deployment
Dockerized and deployed on Hugging Face Spaces.
Input coverage
PDF/DOCX-based course materials and exam papers, with OCR support for scanned documents.
Generation safeguards
Source citations, marks-aware answer checks and automatic retry on quality-gate failure.
Output
Paginated, print-ready PDF study packs.
Preparation modes
Study Guide Pack + Solved Answer Pack.
Evaluation
The current system is evaluated primarily through deterministic pipeline checks and scenario testing rather than a formal benchmark.

What I'd change

Next steps

  • Build a labeled evaluation set covering question parsing, retrieval relevance, citation correctness and answer quality.
  • Replace TF-IDF-only retrieval with hybrid lexical + semantic retrieval once the corpus becomes larger.
  • Add document-level and question-level caching to reduce repeated model calls and API overhead.
  • Expand multimodal ingestion beyond PDFs to support visual content in DOCX and PPTX files.
  • Add structured observability for latency, token usage, retrieval quality and retry frequency.
  • Add automated regression tests so parser or prompt changes can be evaluated against the same benchmark set.