Three generic variants · 7 October 2026 · full test run
Three generic variants of the same résumé (one column on two pages, one column on one page, two columns on one page) were built fresh and read back by thirteen text-extraction engines, simulating the process that an applicant tracking system (ATS) uses before parsing a CV. All three variants pass every test: every engine expected to handle the layout returns every phrase, in order, with ordinary spaces between words. The only shortfalls come from engines with known limitations: three that cannot separate two columns lose about a third of the two-column variant’s phrases, and OCR misreads 6–13 phrases per variant.
All three variants present the same Cybersecurity & AI Researcher résumé.
The number before c is the number of columns and the number before
p the number of pages, so 1c1p is one column on one page.
The one-page variants are shortened versions of the full two-page text. In
2c1p the left column holds About, Experience, Education and Projects,
and the right column holds the remaining sections.
| 1c2p | 1c1p | 2c1p | |
|---|---|---|---|
| columns × pages | 1 × 2 | 1 × 1 | 2 × 1 |
| compressed style | no | yes | yes |
| other layout changes | none | margins 1.2/1.1 cm, name 32 pt, tagline 13 pt, section 12 pt, tighter gaps | name 34 pt, tagline 14 pt, tighter header gaps |
| photo | 4.0 cm | none | 3.7 cm |
| PDF size | 2,755,528 bytes | 152,721 bytes | 2,747,921 bytes |
| phrases checked (body + header) | 91 (83 + 8) | 76 (68 + 8) | 73 (65 + 8) |
A phrase is one piece of text that must survive extraction intact: a heading, an entry, a bullet, a contact field. Each variant is checked against its own phrases.
Before an ATS can recognize a name, a job title or a date, it has to pull the text out of the PDF, and extraction software differs in how it does that. Every PDF is therefore read by thirteen engines: the libraries behind common PDF viewers and Linux tools (poppler, MuPDF, PDFium, pdf.js, Ghostscript), Python and Java document pipelines (pdfminer, PDFBox, Tika), several in more than one mode, and Tesseract, an OCR engine that reads the rendered page like a scanner.
Each engine is either strict or advisory for a given layout.
Tesseract is advisory everywhere: it reads pixels, not the PDF’s text, so its errors measure the OCR rather than the document. In two columns, the engines that rebuild lines across the full page width merge the columns, so they are advisory there, and two engines that sort text by position are advisory for reading order only. In one column, every engine except Tesseract is strict.
| Engines | 1c2p | 1c1p | 2c1p |
|---|---|---|---|
| tesseract | advisory | advisory | advisory |
| poppler-layout, ghostscript, pdfbox-sorted | strict | strict | advisory |
| poppler, pdfminer | strict | strict | strict for phrases, advisory for order |
| the other seven | strict | strict | strict |
Five tests can fail a résumé; the word-gap test only reports. Every variant was built and tested from scratch.
| Test | 1c2p | 1c1p | 2c1p |
|---|---|---|---|
| real word spaces | pass | pass | pass |
| word-gap characters | report clean | report clean | report clean |
| icons as whitespace | pass | pass | pass |
| PDF/A-2B conformance | pass | pass | pass |
| text extraction | pass | pass | pass 3 advisory losses |
| reading order | pass | pass | pass 4 advisory < 1.00 |
What it tests. Whether each gap between words is stored in the PDF as an actual space character, not just as extra distance between letters. Without one, an extractor has to guess where words end, and some guess wrong and run words together. The table counts space characters per text font.
| Font | 1c2p | 1c1p | 2c1p |
|---|---|---|---|
| Inter-Light (body text) | 621 | 399 | 422 |
| Inter-LightItalic | 51 | 46 | 42 |
| Inter-SemiBold | 11 | 11 | 9 |
| SourceSans3-Light | 35 | 77 | 56 |
| SourceSans3-Semibold | 52 | 24 | 23 |
| total, 5 text fonts | 770 | 557 | 552 |
| result | pass | pass | pass |
Discussion. All three variants store every word gap as a real space character, so no extractor has to guess where words end. The totals only follow the amount of text. Neither the compressed style nor two columns changes how gaps are stored, since that is set by the typesetting engine and the template, not by the layout.
What it tests. Test 1 checks how word gaps are stored in the PDF; this one checks what each engine returns for them. It lists any gap that comes back as something other than an ordinary space, above all the no-break space, which Java-based parsers do not treat as a separator and so read the two words as one. It only reports and never fails.
| 1c2p | 1c1p | 2c1p | |
|---|---|---|---|
| engines with only U+0020 gaps | 13/13 | 13/13 | 13/13 |
| gaps that would join words in Java | 0 | 0 | 0 |
| result | report clean | report clean | report clean |
Discussion. All thirteen engines, Tesseract included, return a plain space for every word gap in all three variants, consistent with test 1: a real space character leaves nothing an engine could turn into a no-break space. No Java-based parser would run two words together.
What it tests. The contact line marks each field with a small icon (an envelope, the LinkedIn logo). Every icon must extract as a blank space, not as a stray letter or icon name that would corrupt the email or link next to it.
| Icon font | 1c2p | 1c1p | 2c1p |
|---|---|---|---|
| FontAwesome5Brands-Regular | 2 ok | 2 ok | 2 ok |
| FontAwesome5Free-Solid | 4 ok | 4 ok | 4 ok |
| result | pass | pass | pass |
Discussion. All three variants share the same contact block (two brand icons for LinkedIn and GitHub, four solid icons for the other fields), and all six icons extract as a plain space in every build. This is expected: icon extraction is set by the template, so layout choices cannot change it.
What it tests. PDF/A is the archival PDF standard: every font embedded, no outside dependencies, a file that opens the same everywhere. The reference validator veraPDF confirms that each file really conforms to PDF/A-2B, since the PDF/A label in the metadata proves nothing on its own.
| Check | 1c2p | 1c1p | 2c1p |
|---|---|---|---|
| identification | PDF/A-2B | PDF/A-2B | PDF/A-2B |
| fonts embedded | all | all | all |
| output intent | present | present | present |
| embedded files / forbidden actions / encryption | none | none | none |
| 1.7, tagged | 1.7, tagged | 1.7, tagged | |
| veraPDF 1.30.2 | pass | pass | pass |
Discussion. All three files are valid PDF/A-2B. The photo, about
2.6 MB of 1c2p and 2c1p, does not affect conformance,
but it is the whole size difference, so the photo-less 1c1p suits
portals that limit upload size.
What it tests. The core test: every phrase must appear, complete, in each engine’s text. Differences that do not change the meaning (dash and quote styles, ligatures, letter case) are ignored.
| Engine | 1c2p | 1c1p | 2c1p |
|---|---|---|---|
| poppler | 91/91 | 76/76 | 73/73 |
| poppler-raw | 91/91 | 76/76 | 73/73 |
| poppler-layout | 91/91 | 76/76 | 49/73 advisory |
| mupdf | 91/91 | 76/76 | 73/73 |
| pdfium | 91/91 | 76/76 | 73/73 |
| pdfminer | 91/91 | 76/76 | 73/73 |
| ghostscript | 91/91 | 76/76 | 50/73 advisory |
| pdfjs-old | 91/91 | 76/76 | 73/73 |
| pdfjs-new | 91/91 | 76/76 | 73/73 |
| pdfbox | 91/91 | 76/76 | 73/73 |
| pdfbox-sorted | 91/91 | 76/76 | 49/73 advisory |
| tika | 91/91 | 76/76 | 73/73 |
| tesseract (OCR) | 78/91 advisory | 70/76 advisory | 65/73 advisory |
The three engines that lose phrases in 2c1p rebuild each printed
line across the full page width, so a phrase that wraps in one column gets the
other column’s text inserted where it breaks. Accordingly, 21–23 of the
missing phrases are long, wrapping ones, spread over every section.
poppler-layout extracts from 2c1p, verbatim• Built a federated learning anomaly-detection and classification Personal: Data-driven, Independent, PoC for payment fraud detection, uncovering undetected fraud Organized, Systematic, Adaptable and cutting detection time from weeks to minutes.
An Experience bullet (left column) and the Personal skills line
(right column) merged line by line; ghostscript and
pdfbox-sorted do the same.
| Engine (2c1p) | Lost | By section |
|---|---|---|
| poppler-layout | 24 | experience 7, skills 4, education 3, volunteering 3, projects 2, courses 2, publications 2, about 1 |
| ghostscript | 23 | experience 5, education 4, skills 4, volunteering 3, projects 2, courses 2, publications 2, about 1 |
| pdfbox-sorted | 24 | experience 7, skills 4, education 3, volunteering 3, projects 2, courses 2, publications 2, about 1 |
Discussion: two columns. Every strict engine returns every phrase in all
three variants, so the text stored in the PDF is complete whatever the layout.
The only difference comes from two columns: in 2c1p,
poppler-layout, ghostscript and pdfbox-sorted
recover only 67–68% of the phrases. This is not a fault in the PDF: these
engines cannot separate side-by-side columns, and they return 100% of both
one-column variants. The compressed style causes no losses.
| Cause | 1c2p | 1c1p | 2c1p |
|---|---|---|---|
capital AI read as Al | 6 | 2 | 5 |
diacritics dropped (Chudá, Kučera) | 4 | 3 | 2 |
Scholar URL misread (9_uI59…) | 1 | 1 | 1 |
stray character (model → modell, Goldschmidt → Goldschmiadt) | 2 | 0 | 0 |
| total | 13/91 | 6/76 | 8/73 |
Discussion: Tesseract OCR. Tesseract reads a rendered image of the page,
not the text stored in the PDF, so its misses are OCR character errors rather
than layout problems. As the table shows, they concentrate on the capital
AI, which the sans-serif face makes hard to tell from Al,
and on Czech and Slovak diacritics. Their number follows how often these strings
occur, so the shortest variant, 1c1p, has the fewest. One stray
character is new since the previous report: the 1c2p surname, set in
Source Sans 3 instead of Source Sans Pro, comes back as
Goldschmiadt. Rendering at 600 dpi does not help: the
AI, diacritic and URL errors recur unchanged, and a few different
stray characters replace the old ones, for 13 / 9 / 9 misses against 13 / 6 / 8
at 300 dpi.
Possible fix. The AI errors, and the Scholar URL, whose
capital I comes back as l, have one cause: in Inter, the
capital I and the lowercase l are both a plain vertical stroke, so OCR has to
guess. The fix is a body typeface whose capital I and lowercase l differ in their
regular shapes, for example an I with serifs or an l with a tail. The dropped
diacritics are a different matter: they come from OCR running with an English-only
language model, as in this test setup, and no change of font would fix them.
Practical relevance. In practice, the OCR errors matter little. ATS parsers work primarily with the text stored in the PDF, and since this file is tagged, PDF/A-2B-compliant and its text reads back completely, a fallback to pure OCR is unlikely. The OCR results are therefore for informational purposes only.
What it tests. Finding every phrase is not enough if they come back scrambled: an ATS that reads a date after the wrong employer files it under the wrong job. The test scores how much of each engine’s output keeps the order the text is stored in, from 0 to 1 (perfect). The body must score 1.00 in every strict engine; the contact header is only reported, since the order of contact fields carries no meaning.
| Engine — body order score | 1c2p | 1c1p | 2c1p |
|---|---|---|---|
| poppler | 1.00 (83/83) | 1.00 (68/68) | 0.86 (56/65) advisory |
| poppler-raw | 1.00 (83/83) | 1.00 (68/68) | 1.00 (65/65) |
| poppler-layout | 1.00 (83/83) | 1.00 (68/68) | 0.71 (29/41) advisory |
| mupdf | 1.00 (83/83) | 1.00 (68/68) | 1.00 (65/65) |
| pdfium | 1.00 (83/83) | 1.00 (68/68) | 1.00 (65/65) |
| pdfminer | 1.00 (83/83) | 1.00 (68/68) | 1.00 (65/65) advisory |
| ghostscript | 1.00 (83/83) | 1.00 (68/68) | 0.71 (30/42) advisory |
| pdfjs-old | 1.00 (83/83) | 1.00 (68/68) | 1.00 (65/65) |
| pdfjs-new | 1.00 (83/83) | 1.00 (68/68) | 1.00 (65/65) |
| pdfbox | 1.00 (83/83) | 1.00 (68/68) | 1.00 (65/65) |
| pdfbox-sorted | 1.00 (83/83) | 1.00 (68/68) | 0.71 (29/41) advisory |
| tika | 1.00 (83/83) | 1.00 (68/68) | 1.00 (65/65) |
| tesseract (OCR) | 1.00 (73/73) | 1.00 (64/64) | 1.00 (59/59) |
| header, all engines | 1.00 | 1.00 except poppler 0.88 (7/8) report | 1.00 |
In brackets: phrases in order / phrases found, so an engine that lost
phrases is scored on fewer of them. The reference order comes from
poppler-raw, which outputs the text exactly as stored.
linkedin.com/in/… from the end of the first contact line to after the
Scholar link on the second. The field is complete, only later in the output.poppler-raw): section headings in 2c1pAboutExperienceEducationProjectsSkillsCourses & CertificationsSelected PublicationsVolunteering
The left column first, then the right one, as the PDF stores them.
poppler extracts from 2c1p, verbatim (abridged)About Skills Researcher with 9 years of applied research and R&D experience,… (rest of the About paragraph) Professional: Artificial Intelligence &…Tools: Python for AI (numpy, pandas,…Personal: Data-driven, Independent,…Languages: English (full professional…Experience
The right column’s Skills heading and block (highlighted) come out next to About, level with it at the top of the page, so an ATS would see Skills before Experience. These are all nine of poppler’s out-of-order phrases; every other section keeps its order.
Discussion. In both one-column variants every engine returns the body in
document order, so order is a problem only in two columns. There, the seven
engines that output text as stored (poppler-raw, mupdf,
pdfium, both pdf.js builds, pdfbox,
tika) keep a perfect 1.00, because the template stores each column as
one continuous block. The engines that re-sort text by position depart from it to
different degrees: pdfminer still reaches 1.00, poppler
moves one block (0.86), and the three side-by-side engines also scramble what
they recover (0.71).
| Test | Passes in all three? | Where variants differ |
|---|---|---|
| real word spaces | yes | no difference: 770 / 557 / 552 real space characters |
| word-gap characters | report | no difference: U+0020 only, 13 of 13 engines |
| icons as whitespace | yes | no difference: 6 of 6 icons in each |
| PDF/A-2B | yes | no difference |
| text extraction | yes | 2c1p: 3 side-by-side engines lose 23–24 phrases (advisory); Tesseract misses 13 / 6 / 8 everywhere (advisory) |
| reading order | yes | 2c1p: 4 position-sorting engines at 0.71–0.86 (advisory); 1c1p: poppler moves LinkedIn in the header (reported) |
poppler also reorder the body. An ATS using one of them would receive
merged columns, yet the PDF is sound: the seven engines that keep the stored order
read it perfectly.poppler-raw, which therefore always scores 1.00). A PDF that stored its
text in the wrong order would shift the reference rather than fail the test.1c1p skills
line) was checked by hand.| lualatex | LuaHBTeX 1.24.0 (TeX Live 2026) |
|---|---|
| python | 3.14.6 (env/python) |
| pymupdf | 1.28.2 |
| pypdfium2 | 5.13.0 |
| pdfminer.six | 20260107 |
| poppler | pdftotext 26.01.0 |
| ghostscript | 10.06.0 |
| tesseract | 5.5.3 |
| java | OpenJDK 25.0.3 |
| pdfbox / tika | 3.0.5 / 3.2.3 |
| node | v20.18.1 |
| pdf.js | 2.5.207 (old), 4.10.38 (new) |
| veraPDF | 1.30.2 |
From tests/, one command per variant:
| 1c2p | ./pipeline.sh ../realizations/generic_research-1c2p/main.tex |
|---|---|
| 1c1p | ./pipeline.sh ../realizations/generic_research-1c1p/main.tex |
| 2c1p | ./pipeline.sh ../realizations/generic_research-2c1p/main.tex |
Each run generates phrases.txt, builds
<label>.pdf, runs run.sh and writes
run.log and pipeline.json to
tests/out/generic_research-<variant>/. The 600 dpi
pass renders each built PDF with pdftoppm -r 600 -gray -png, runs
tesseract on every page, and scores the output with the same
check.py and order.py.
1c1p untested, is now checked too.All three pipeline runs exited 0. Every figure here comes from their
run.log, pipeline.json, order.json and the 13
per-engine extractions. The 2c1p loss breakdown and the Tesseract diagnosis were
computed from those files with the tests’ own normalization.