Project

Handwriting Synthesis Pipeline for Dataset Generation

December 2024

Academic View on GitHub

Overview

Handwritten student solutions are one of the most annoying input formats for OCR. You get free form handwriting, math notation, mixed languages, messy corrections, and scan artifacts, all in the same page. This project builds a data generation pipeline that produces that kind of input on purpose.

The core idea is simple: generate math exercises in LaTeX, render them in a handwriting style, then degrade the result until it looks like a realistic scan. The output is a large set of PNG images that can be used to stress test OCR systems and to train models that need more coverage than real data usually provides.

Example of a generated handwritten solution

Context

This project was conducted for ETH Zurich and connects to Ethel, ETH’s virtual teaching assistant. The practical problem behind it is straightforward: if you want a teaching assistant to read student submissions, you need models that can handle handwritten math and multilingual text under real world noise. That requires data that actually looks like student work, including mistakes and corrections, not just clean synthetic renders.

My role

I co led dataset and pipeline design and worked closely with partners from the Swiss AI Center and ETH Zurich. My focus was on turning the initial concept into something you can run end to end: document generation, handwriting simulation, mistake rendering, and image level augmentation.

What We built

The pipeline produces synthetic handwritten math solutions with three properties that matter for OCR training.

First, variety in content. The generator creates full solutions, not just isolated formulas, including explanatory text and intermediate steps.

Second, variety in appearance. Different handwriting fonts, spacing irregularities, page styles, and rendering settings produce visually distinct samples.

Third, realism through imperfection. The pipeline injects mistakes that get crossed out, and it adds scan like artifacts such as blur and noise.

How the pipeline works

  1. LaTeX generation
    A language model generates math exercises and their solutions as LaTeX. The output is cleaned so it reliably forms a compilable document. Language is sampled from a predefined set, which makes the dataset multilingual by construction.

  2. Mistake simulation with visible corrections
    To mimic student corrections, the generator inserts intentional mistakes and marks them with a dedicated LaTeX command. During compilation, the header randomly selects one of several strike through styles implemented in TikZ. This gives you different “human looking” correction patterns rather than one uniform line.

Example of a correction rendered as a strike through

  1. Handwriting simulation for text and math
    Standard handwriting fonts can look good for text but usually fail once you render mathematical symbols. The project uses handwriting fonts where possible and complements them with a custom font pipeline for math heavy content. On top of that, a preprocessing command introduces small word level shifts and rotations so lines do not look typeset perfect.

Example of a formula rendered with the custom font

  1. Rendering to images
    Each LaTeX document is compiled with XeLaTeX. The pipeline can generate multiple visual variants from the same source by changing fonts, text colors, page colors, and optional grid paper backgrounds. PDFs are then converted into high resolution PNG images.

  2. Image degradation
    To make the samples resemble scans, the pipeline produces augmented variants of each image. Typical outputs include a clean render, a noisy render, a blurred render, and a noisy plus blurred render. This is the step that usually makes synthetic data stop looking “too clean”.

Blurred and noisy render
Blurred + Noisy
Blurred render
Blurred

Custom font generation for math content

A key subproject is the custom font generation pipeline. The workflow is built around collecting glyphs via templates, extracting each character, converting it into scalable vectors, and assembling everything into an OpenType font.

This matters because math OCR is fragile when symbols look out of distribution. If your dataset only supports text handwriting and falls back to a different style for math symbols, you end up training models on inconsistencies that do not occur in real student submissions. With a math capable handwriting font, equations and surrounding text look like they belong to the same pen.

Results

The output is a dataset of realistic handwritten math solutions designed for OCR training and evaluation. It covers multilingual text, structured notation, layout variation, corrections, and scan artifacts. In practice, that combination is what tends to break OCR systems first, so it is exactly what you want when you are stress testing models or expanding a training set.

The project received an “Excellent” evaluation.

What I would improve next

The pipeline already produces realistic single page solutions, but there is still headroom in document diversity. The next improvements I would prioritize are richer layouts, marginal notes, small diagrams, partial scans, and other artifacts like stains or uneven illumination. Another obvious extension is blackboard style data, since lecture content often comes from photos rather than scans.

On the generation side, a stroke based handwriting synthesizer for formulas would remove some remaining font constraints and make symbol placement even less uniform.

The repository is linked at the top of this page.