Document processing
A document is usable for question generation only after text extraction finishes and a structure tree exists.
Files you can upload
Section titled “Files you can upload”| Kind | Extensions | How the text is obtained |
|---|---|---|
.pdf | Embedded text is read first. Pages without usable text go through optical character recognition | |
| Word | .docx | Text is read from the file |
| PowerPoint | .pptx | Text is read from the slides |
| Plain text | .txt | Read as-is |
| Image | .png, .jpg, .jpeg | Optical character recognition |
A scanned PDF is read by character recognition, not by sending page images to a chat model. Recognition is set up for Urdu, Arabic, and English unless your deployment has more language packs installed. Changing the teacher app’s interface language does not install a pack and does not re-read a file.
Two statuses, not one
Section titled “Two statuses, not one”Text extraction and structure extraction are tracked separately. A document can have finished one and not the other.
| Text extraction | Meaning |
|---|---|
none | Not started |
queued | Waiting for a worker |
processing | Being read now |
done | Text is available |
failed | Could not be read |
| Structure extraction | Meaning |
|---|---|
none | Not started |
queued | Waiting |
processing | Being analysed |
ready | A structure tree is available |
failed | Could not be produced |
Large documents are placed on a slower lane so one textbook cannot stall every short file in the workspace. Queued is a wait, not a failure.
The structure tree
Section titled “The structure tree”Extraction proposes chapters, sections, and topics. Each entry can have a summary and a page range. You can edit the title, summary, and page range. An entry you change is marked edited, so you can see what the extractor proposed and what you corrected.
Regenerating the tree snapshots the previous one, so an earlier version can be restored.
Question generation uses the entries you select, not the whole file by default. A weak tree produces weak questions. Fix the tree before you spend credits on a draft.
Search by meaning
Section titled “Search by meaning”After text extraction, the document is split into passages and indexed so ZeroExam can find a paragraph by meaning rather than by the exact words. That index is what grounds question generation and the assistant in your material. It is not a separate product you have to turn on.
Organising files
Section titled “Organising files”Documents can group files into folders and attach them to a subject. Moving a file does not re-run extraction. Duplicating a file copies it; it does not share processing state with the original.