TL;DR. Models learn new documents by training on text a model writes about them, and that text can be no better than its writer and always loses some of the page. We propose the sketch: the page itself, with the words worth remembering hidden and the reason each one matters. The learner fills the page back in, so everything it learns comes from the source, and a data engine learns, round by round, to hide what the learner is missing.
Motivation: every way of studying a document is a "view" of it
Models learn a new document by training on text written about it. Synthetic text already shapes pretraining (Eldan and Li, 2023; Gunasekar et al., 2024; Maini et al., 2024) and instruction tuning (Wang et al., 2023). For new knowledge it takes the form of implications (Akyürek et al., 2024; Lampinen et al., 2025), QA pairs (Park et al., 2025), entity-graph essays (Yang et al., 2025) and self-chosen study strategies (Lin et al., 2025). The reason it helps is that a model extracts a fact reliably only when training presents it in varied forms (Allen-Zhu and Li, 2024). SEAL (Zweiger et al., 2025) makes the writing a learning problem. A model writes study material for a passage, a copy of the model trains on it and is quizzed, and the material that taught most trains the writer.
We call each of these a "view" of the document: a presentation of the same content that the learner trains on with next-token loss. A generated view has two problems built in.
(i) Capped by the writer. A model can only write about a passage what it understands of it, and its mistakes become training targets. When the writer is the learner, the model is asked to teach itself what it does not yet know, using only what it already knows.
(ii) Inherent information loss. A rewrite drops details, an implication adds claims the page never made. Here is the implications view the base model wrote for a SQuAD passage on Genghis Khan's campaign against Khwarezmia:
1. Internal Instability: The Shah's empire was plagued by serious internal division
("diverse internecine feuds"), which critically weakened his military resistance.
2. Poor Military Strategy: The Shah's decision to fragment his army into isolated groups
made his forces vulnerable to targeted attacks.
3. Weakness in Leadership/Courage: The Shah was unwilling to fight to the end and
preferred to flee rather than accept defeat.
...
"Weakness in Leadership/Courage" is interpretation the page never states, and the learner trains on it as fact. Across 50 passages, only 7% of the word trigrams in implications appear in the passage; 10% for rewrites, 36% for QA pairs.
What we expect: generated views hit a ceiling, the page does not
Give a learner more of a generated view and it should help, then flatten: more text from the same writer repeats what the writer understood and adds more of what the view got wrong. The limit comes from the writer.

What we aim for
We want a view that is grounded, so everything the learner is trained to produce comes from the document; scalable, so it gives more accuracy per token trained and keeps rising where generated views flatten; and learnable, so a data engine improves it from the learner's own results, without a stronger teacher.
The sketch: the page, with the right words hidden
A sketch is what a good student does with a highlighter. For the same passage, the data engine writes:
HIGHLIGHT diverse internecine feuds
WHY What caused the army to be split?
HIGHLIGHT Mongol army
WHY What was able to quickly seize Otrar?
HIGHLIGHT Inalchuq
WHY Who was executed?
...
A HIGHLIGHT is words copied from the page; code drops any it cannot find there. A WHY is the question those words answer. Together they form one record, a StudyItem: claims, each with its page words, where they occur, and its question. At most 15% of the passage is hidden.
The learner sees the page with each claim hidden in place and its question listed after it:
Some words of the passage below are hidden and marked [[1]], [[2]], and so on ...
[[1]]
The [[2]] was split by [[3]] and by the [[4]] to divide his army into [[5]] ...
The [[10]] quickly seized the town of Otrar ... [[1]] ordered the wholesale massacre
of many of the civilians ... and executed [[11]] by pouring molten silver into his
ears and eyes ...
What each blank answers:
[[1]] Who is mentioned first?
[[2]] Whose army was split?
[[3]] What caused the army to be split?
...
[[11]] Who was executed?
and fills in the page's words:
[[1]] Genghis Khan
[[2]] Shah's army
[[3]] diverse internecine feuds
...
[[11]] Inalchuq
It is ordinary next-token training, like cloze or fill-in-the-middle, with the whole page around every blank. The loss falls only on the hidden page words, plus a plain read of the passage that every view includes. The WHY is a cue, never a target.
Every token the learner produces is a page token, so nothing is lost between the view and the document. A weak sketcher picks less useful blanks, but every blank it picks is correct. And a sketch is small: about 50 trained tokens per passage, against 300 to 1,000 for generated views. The WHY lets a sketch reorganize knowledge without generating any, by deciding which page words form one fact and which question that fact answers.
Learning to sketch is Feynman studying
A sketch is only as good as the words it hides, and the data engine can learn to choose them. We keep SEAL's training loop (ReST-EM; Singh et al., 2024) and change only what the data engine writes. Each round is one Feynman study step:
| Feynman step | in each round |
|---|---|
| mark what to learn | the data engine sketches training passages, several samples each |
| test yourself | a fresh learner studies each sketch, then answers the passage's questions closed book |
| study the gaps | the sketch that taught most per passage is kept, and the data engine is fine-tuned to write it |
A blank on something the learner already knew teaches nothing, so it earns
nothing, and round over round the data engine should learn to hide what
the learner is missing. The data engine only learns to point at page
words, so it can never learn to inject content. To track this, we ask the
model to explain each topic from memory, from the title alone, and report
known_rate: the share of a sketch's blanks that explanation already
states. It should fall as the rounds go on.
Experimental framework
We follow SEAL's setup and change only the view. The data are SEAL's SQuAD (Rajpurkar et al., 2016) passages and selection; each training round uses new passages, and evaluation uses validation passages the data engine never trained on. All models are open: Gemma 4 E4B-it is both the data engine and the learner, one set of weights as in SEAL, and Gemma 4 12B-it grades the answers.
We compare the passage alone; SEAL's implications (with its long,
very-long and chain-of-thought variants), rewrite and self-qa;
EntiGraph; sketch@random, which hides random words; and sketch@recall,
our sketch. For learning to write views, SEAL's implications and
sketch@recall each go through their own ReST-EM rounds. We measure
closed-book accuracy two ways, as SEAL does: one learner per passage, and
one learner that studies the whole corpus.
The pilot below uses 50 validation passages and one small round. The full run, at SEAL's scale (200 validation passages, 50 training passages per round, three rounds), and the scaling curves are in progress.
Initial signal

| view | single passage | corpus | tokens added per passage |
|---|---|---|---|
| rewrite | 72 | 59.2 | 483 |
| implications | 66 | 56.3 | 346 |
| self-qa | 64 | 56.3 | 315 |
| implications-very-long | 63 | 46.6 | 1,053 |
| sketch@recall | 60 | 52.9 | 52 |
| the passage alone | 59 | 45.0 | 0 |
| sketch@random | 54 | 45.8 | 62 |
Closed-book accuracy (%); the model answers 25% before studying.
What is hidden matters. At the same budget, sketch@recall beats random
masking by 6 points per passage and 7 on the corpus, and random masking
barely beats re-reading.
The sketch is the most token-efficient view, adding 3.0 corpus points per thousand tokens trained against 0.6 to 0.7 for generated views. Generated views already show a ceiling: the longer implications formats train on 2 to 3 times more tokens and score 7 to 10 points lower on the corpus.
Accuracy still trails the best generated views. The knowledge gets in, since the sketch raises the gold answer's likelihood more than any generated view, but the learner often answers in the fill-in's style, as a list of page phrases. That is what we are fixing first.

After one round, sketch@recall is the better corpus learner of the two
arms (55.5 against 48.7).
Next steps and limitations
- Scaling curves, running now: every view at 1, 2, 5, 10 and 20 samples per passage on 200 validation passages, with a fitted ceiling for each. This is the direct test of Figure 1.
- The full run, also running: three ReST-EM rounds per arm at SEAL's scale. The pilot's single small round is noisy, and its 24-example fine-tuning step visibly shifts the learner.
- Practice at answering. We tried having the learner also answer each WHY closed book, with the page words as the only target. It backfired: the learner memorized a few short phrases and gave them to any question on the topic, and accuracy fell from 60% to 36%. Closing the gap between likelihood and answers needs a form of practice that does not collapse onto the WHY lines.
- A better use of the budget. The sketch keeps claims in reading order, so the budget can run out before a passage's key facts, and selecting by accuracy made the pilot's sketches shrink each round. The full run selects only full-budget sketches, and we are testing a prompt that asks for the most important facts first.
- Model size. With Gemma 4 E2B, E4B and 12B, generated views should degrade faster than the sketch as the model shrinks, if the writer's ceiling is real.
- Harder corpora. SQuAD asks for stated facts, and almost everything a sketch hides is new to the model. Corpora the model partly knows will test Feynman-style skipping, and multi-hop sets like MuSiQue will test whether WHY lines can connect facts across passages. The pilot is also one seed, with a 12B judge standing in for GPT-4.1.
Related work
SEAL (Zweiger et al., 2025) is the loop we build on, and its formats are our baselines. EntiGraph (Yang et al., 2025) and Active Reading (Lin et al., 2025) generate diverse text about a corpus; the sketch generates none. Masked language modeling (Devlin et al., 2019), span corruption (Raffel et al., 2020) and fill-in-the-middle (Bavarian et al., 2022) train on hidden spans of real text; here the spans are chosen, with a reason, by a model that learns to choose them. ReST-EM (Singh et al., 2024) is the training algorithm.
Citation.
@misc{sketch2026,
title = {Learning to Sketch},
author = {Yoon, Lauren Hyoseo},
year = {2026},
note = {Blog post}
}
References
Akyürek, A. F., Akyürek, E., Choshen, L., Wijaya, D., and Andreas, J. (2024). Deductive closure training of language models for coherence, accuracy, and updatability. Findings of ACL 2024. https://aclanthology.org/2024.findings-acl.584/
Allen-Zhu, Z. and Li, Y. (2024). Physics of language models: Part 3.1, knowledge storage and extraction. https://arxiv.org/abs/2309.14316
Bavarian, M., Jun, H., Tezak, N., Schulman, J., McLeavey, C., Tworek, J., and Chen, M. (2022). Efficient training of language models to fill in the middle. https://arxiv.org/abs/2207.14255
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. NAACL 2019. https://aclanthology.org/N19-1423/
Eldan, R. and Li, Y. (2023). TinyStories: How small can language models be and still speak coherent English? https://arxiv.org/abs/2305.07759
Gunasekar, S., Zhang, Y., Aneja, J., et al. (2024). Textbooks are all you need. https://openreview.net/forum?id=Fq8tKtjACC
Lampinen, A. K., Chaudhry, A., Chan, S. C. Y., et al. (2025). On the generalization of language models from in-context learning and finetuning: a controlled study. https://arxiv.org/abs/2505.00661
Lin, J., Berges, V.-P., Chen, X., Yih, W.-t., Ghosh, G., and Oğuz, B. (2025). Learning facts at scale with Active Reading. https://arxiv.org/abs/2508.09494
Maini, P., Seto, S., Bai, R., Grangier, D., Zhang, Y., and Jaitly, N. (2024). Rephrasing the web: A recipe for compute and data-efficient language modeling. ACL 2024. https://aclanthology.org/2024.acl-long.757/
Park, C. F., Zhang, Z., and Tanaka, H. (2025). New News: System-2 fine-tuning for robust integration of new knowledge. https://arxiv.org/abs/2505.01812
Raffel, C., Shazeer, N., Roberts, A., et al. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR 21(140). https://jmlr.org/papers/v21/20-074.html
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. (2016). SQuAD: 100,000+ questions for machine comprehension of text. EMNLP 2016. https://aclanthology.org/D16-1264/
Singh, A., Co-Reyes, J. D., Agarwal, R., et al. (2024). Beyond human data: Scaling self-training for problem-solving with language models. TMLR. https://openreview.net/forum?id=lNAyUngGFK
Wang, Y., Kordi, Y., Mishra, S., et al. (2023). Self-Instruct: Aligning language models with self-generated instructions. ACL 2023. https://aclanthology.org/2023.acl-long.754/
Yang, Z., Band, N., Li, S., Candès, E., and Hashimoto, T. (2025). Synthetic continued pretraining. ICLR 2025. https://openreview.net/forum?id=07yvxWDSla
Zweiger, A., Pari, J., Guo, H., Akyürek, E., Kim, Y., and Agrawal, P. (2025). Self-adapting language models. NeurIPS 2025. https://arxiv.org/abs/2506.10943