Teaching a Small Hebrew Model to Explain Bureaucracy: Fine-Tuning DictaLM 1.7B with Teacher-Generated Data
Ahmad Tawil · Sep 2026
The question
Can a small, open Hebrew model that runs on a laptop learn to answer in plain Hebrew, cite sources and say “I don't know”?
01
TL;DR
- 1
Generated a training set from 287 Kol Zchut articles with a 12B teacher (vLLM), then fine-tuned DictaLM 3.0 1.7B with LoRA (Unsloth + TRL) in about 20 minutes on one GPU.
- 2
On 232 held-out examples from articles never seen in training, correct refusals went from 0% to 61%, letter-format compliance from 28% to 100%, and correct answers doubled.
- 3
The 12B teacher never refused. Refusal is a behaviour the data design taught the student, not one it copied, at the cost of 12% over-refusal.
02
Key results
- correct refusals
- 0% → 61%correct refusals
- letter-format compliance
- 28% → 100%letter-format compliance
- correct answers (manual)
- 8 → 16 / 30correct answers (manual)
- per answer on a laptop CPU
- 51 → 33 sper answer on a laptop CPU
Correct refusals when sources don't answer
Letter summary in the required format
Plain text (no markdown)
Correct answers (manual, 30 questions)
03
In plain words
People who lose a job or receive an official letter often don't know what they're entitled to. Kol Zchut has the answers, but in long, formal Hebrew, and a chatbot on a commercial LLM would send personal documents to an outside service.
This paper asks whether a 1.7B-parameter open model can be taught four behaviours: grounded answers with [n] citations, a fixed refusal when the sources don't cover the question, plain-Hebrew rewrites, and summaries of official letters under four fixed headings.
The training data was generated by a 12B teacher and filtered heavily; unanswerable examples were checked with the teacher as a judge, so only 148 of 1,500 candidates were kept. The split is by article, with context leakage removed.
Honest limitations
- !Teacher-generated data can carry teacher errors.
- !12% over-refusal on answerable questions.
- !Covers four topic areas only.
- !Manual scoring was done by an AI assistant, not a human expert.
- 1
Collect
287 Kol Zchut articles → 1,688 chunks, via the MediaWiki API
- 2
Generate
DictaLM 3.0 12B teacher on vLLM: grounded QA, hard negatives, letters, rewrites
- 3
Filter
citation, format and language checks, near-duplicate removal, manual review
- 4
Fine-tune
LoRA r = 16 on DictaLM 1.7B (Unsloth + TRL), 2 epochs, loss on answers only
- 5
Evaluate
232 held-out examples, blind manual scoring, end-to-end PDF and photo letter tests
04
Read the full paper
05
Cite this work
@misc{tawil2026hebrew,
title = {Teaching a Small Hebrew Model to Explain Bureaucracy: Fine-Tuning DictaLM 1.7B with Teacher-Generated Data},
author = {Tawil, Ahmad},
year = {2026},
month = sep,
url = {https://ahmadtawil.dev/research/hebrew-bureaucracy-llm}
}06