01 / 05 · AI / Education

GradeMIND

Through-lineThe model reads, arithmetic decides.

Category
AI / Education
Role
Builder
Status
Hackathon build · CBSE-grade next

01 / 07Context

01Context

GradeMIND is a handwritten answer-sheet grader, built for a national hackathon.

Marking is a trust problem. A mark has to be explainable and repeatable, and a language model's numeric output is neither.

02Approach

The model reads. Arithmetic decides. Every mark traces to a criterion ID, an evidence span and deterministic arithmetic, and never to an LLM numeric output.

This is the same deterministic core and LLM language layer split used across PRYSM and AHAL AI, applied to grading.

03System

Noisy handwritten scans go through three OCR engines and Gemini Vision. Scores are computed by code.

Diagram · The boundary between reading and deciding
  1. 01

    Three OCR engines

    PaddleOCR, EasyOCR and Tesseract, fused by confidence voting

  2. 02

    Gemini Vision

    Reads noisy handwritten scans

  3. 03

    Question segmentation

    Splits the sheet by question

  4. 04

    ScoreComputer

    Deterministic arithmetic, tied to criterion ID and evidence span

  5. 05

    Annotated PDF

    Every mark traced to an evidence span

Deterministic coreLLM language layer

04Build

  • Three OCR engines fused by confidence voting, with Gemini Vision for noisy handwritten scans
  • Question segmentation and annotated-PDF output tracing every mark to an evidence span
  • Deterministic scorer with reproducibility tests
  • Atomic job-state persistence with resume and cache reuse
  • Human-override preservation

05Challenges

  • Keeping numeric authority out of the model while still using it to read.
  • A production-readiness audit exposed fabricated agent results and a score-desync bug. A six-phase remediation plan now gates real student data behind human review.

06Result

Built for a national hackathon. Next step: CBSE-grade production.

3
OCR engines fused

07Learnings

Verify against real state, not summaries. The audit only mattered because it checked what the agents had actually produced.

AI / SYSTEMS / PRODUCT ENGINEERING
SHREEKUMAR.B000