Paper · Sep 2026Fine-tuningHebrew NLPRAG

Teaching a Small Hebrew Model to Explain Bureaucracy: Fine-Tuning DictaLM 1.7B with Teacher-Generated Data

Ahmad Tawil · Sep 2026

The question

Can a small, open Hebrew model that runs on a laptop learn to answer in plain Hebrew, cite sources and say “I don't know”?

01

TL;DR

  1. 1

    Generated a training set from 287 Kol Zchut articles with a 12B teacher (vLLM), then fine-tuned DictaLM 3.0 1.7B with LoRA (Unsloth + TRL) in about 20 minutes on one GPU.

  2. 2

    On 232 held-out examples from articles never seen in training, correct refusals went from 0% to 61%, letter-format compliance from 28% to 100%, and correct answers doubled.

  3. 3

    The 12B teacher never refused. Refusal is a behaviour the data design taught the student, not one it copied, at the cost of 12% over-refusal.

02

Key results

correct refusals
0% → 61%correct refusals
letter-format compliance
28% → 100%letter-format compliance
correct answers (manual)
8 → 16 / 30correct answers (manual)
per answer on a laptop CPU
51 → 33 sper answer on a laptop CPU
Base DictaLM 1.7B Fine-tuned 1.7B

Correct refusals when sources don't answer

0%
61%

Letter summary in the required format

28%
100%

Plain text (no markdown)

49%
100%

Correct answers (manual, 30 questions)

8/30
16/30
Real results · 232 held-out examples from articles never seen in training

03

In plain words

People who lose a job or receive an official letter often don't know what they're entitled to. Kol Zchut has the answers, but in long, formal Hebrew, and a chatbot on a commercial LLM would send personal documents to an outside service.

This paper asks whether a 1.7B-parameter open model can be taught four behaviours: grounded answers with [n] citations, a fixed refusal when the sources don't cover the question, plain-Hebrew rewrites, and summaries of official letters under four fixed headings.

The training data was generated by a 12B teacher and filtered heavily; unanswerable examples were checked with the teacher as a judge, so only 148 of 1,500 candidates were kept. The split is by article, with context leakage removed.

Honest limitations

  • !Teacher-generated data can carry teacher errors.
  • !12% over-refusal on answerable questions.
  • !Covers four topic areas only.
  • !Manual scoring was done by an AI assistant, not a human expert.
  1. 1

    Collect

    287 Kol Zchut articles → 1,688 chunks, via the MediaWiki API

  2. 2

    Generate

    DictaLM 3.0 12B teacher on vLLM: grounded QA, hard negatives, letters, rewrites

  3. 3

    Filter

    citation, format and language checks, near-duplicate removal, manual review

  4. 4

    Fine-tune

    LoRA r = 16 on DictaLM 1.7B (Unsloth + TRL), 2 epochs, loss on answers only

  5. 5

    Evaluate

    232 held-out examples, blind manual scoring, end-to-end PDF and photo letter tests

04

Read the full paper

Open the PDF on GitHubThe paper is hosted with its code.↗

05

Cite this work

BibTeX
@misc{tawil2026hebrew,
  title  = {Teaching a Small Hebrew Model to Explain Bureaucracy: Fine-Tuning DictaLM 1.7B with Teacher-Generated Data},
  author = {Tawil, Ahmad},
  year   = {2026},
  month  = sep,
  url    = {https://ahmadtawil.dev/research/hebrew-bureaucracy-llm}
}

06

Keep reading