Skip to content
Open app

Document processing

A document is usable for question generation only after text extraction finishes and a structure tree exists.

KindExtensionsHow the text is obtained
PDF.pdfEmbedded text is read first. Pages without usable text go through optical character recognition
Word.docxText is read from the file
PowerPoint.pptxText is read from the slides
Plain text.txtRead as-is
Image.png, .jpg, .jpegOptical character recognition

A scanned PDF is read by character recognition, not by sending page images to a chat model. Recognition is set up for Urdu, Arabic, and English unless your deployment has more language packs installed. Changing the teacher app’s interface language does not install a pack and does not re-read a file.

Text extraction and structure extraction are tracked separately. A document can have finished one and not the other.

Text extractionMeaning
noneNot started
queuedWaiting for a worker
processingBeing read now
doneText is available
failedCould not be read
Structure extractionMeaning
noneNot started
queuedWaiting
processingBeing analysed
readyA structure tree is available
failedCould not be produced

Large documents are placed on a slower lane so one textbook cannot stall every short file in the workspace. Queued is a wait, not a failure.

Extraction proposes chapters, sections, and topics. Each entry can have a summary and a page range. You can edit the title, summary, and page range. An entry you change is marked edited, so you can see what the extractor proposed and what you corrected.

Regenerating the tree snapshots the previous one, so an earlier version can be restored.

Question generation uses the entries you select, not the whole file by default. A weak tree produces weak questions. Fix the tree before you spend credits on a draft.

After text extraction, the document is split into passages and indexed so ZeroExam can find a paragraph by meaning rather than by the exact words. That index is what grounds question generation and the assistant in your material. It is not a separate product you have to turn on.

Documents can group files into folders and attach them to a subject. Moving a file does not re-run extraction. Duplicating a file copies it; it does not share processing state with the original.