Three generic variants · 7 October 2026 · full test run

Lucid CV generic résumé: three variants compared

Three generic variants of the same résumé (one column on two pages, one column on one page, two columns on one page) were built fresh and read back by thirteen text-extraction engines, simulating the process that an applicant tracking system (ATS) uses before parsing a CV. All three variants pass every test: every engine expected to handle the layout returns every phrase, in order, with ordinary spaces between words. The only shortfalls come from engines with known limitations: three that cannot separate two columns lose about a third of the two-column variant’s phrases, and OCR misreads 6–13 phrases per variant.

3 variants13 engines 91 / 76 / 73 phrases15/15 tests passed LuaHBTeX 1.24.0 (TeX Live 2026)

Introduction

The three variants

All three variants present the same Cybersecurity & AI Researcher résumé. The number before c is the number of columns and the number before p the number of pages, so 1c1p is one column on one page. The one-page variants are shortened versions of the full two-page text. In 2c1p the left column holds About, Experience, Education and Projects, and the right column holds the remaining sections.

1c2p1c1p 2c1p
columns × pages1 × 21 × 12 × 1
compressed stylenoyesyes
other layout changesnone margins 1.2/1.1 cm, name 32 pt, tagline 13 pt, section 12 pt, tighter gaps name 34 pt, tagline 14 pt, tighter header gaps
photo4.0 cmnone3.7 cm
PDF size2,755,528 bytes152,721 bytes 2,747,921 bytes
phrases checked (body + header)91 (83 + 8) 76 (68 + 8)73 (65 + 8)

A phrase is one piece of text that must survive extraction intact: a heading, an entry, a bullet, a contact field. Each variant is checked against its own phrases.

How the résumé is read back

Before an ATS can recognize a name, a job title or a date, it has to pull the text out of the PDF, and extraction software differs in how it does that. Every PDF is therefore read by thirteen engines: the libraries behind common PDF viewers and Linux tools (poppler, MuPDF, PDFium, pdf.js, Ghostscript), Python and Java document pipelines (pdfminer, PDFBox, Tika), several in more than one mode, and Tesseract, an OCR engine that reads the rendered page like a scanner.

Strict and advisory engines

Each engine is either strict or advisory for a given layout.

Tesseract is advisory everywhere: it reads pixels, not the PDF’s text, so its errors measure the OCR rather than the document. In two columns, the engines that rebuild lines across the full page width merge the columns, so they are advisory there, and two engines that sort text by position are advisory for reading order only. In one column, every engine except Tesseract is strict.

Engines1c2p1c1p 2c1p
tesseractadvisoryadvisoryadvisory
poppler-layout, ghostscript, pdfbox-sortedstrictstrict advisory
poppler, pdfminerstrictstrict strict for phrases, advisory for order
the other sevenstrictstrictstrict

Scoreboard

Five tests can fail a résumé; the word-gap test only reports. Every variant was built and tested from scratch.

Test1c2p1c1p 2c1p
real word spacespass passpass
word-gap charactersreport clean report cleanreport clean
icons as whitespacepass passpass
PDF/A-2B conformancepass passpass
text extractionpass pass pass 3 advisory losses
reading orderpass pass pass 4 advisory < 1.00

1 · Real word spaces

What it tests. Whether each gap between words is stored in the PDF as an actual space character, not just as extra distance between letters. Without one, an extractor has to guess where words end, and some guess wrong and run words together. The table counts space characters per text font.

Font1c2p 1c1p2c1p
Inter-Light (body text)621399422
Inter-LightItalic514642
Inter-SemiBold11119
SourceSans3-Light357756
SourceSans3-Semibold522423
total, 5 text fonts770 557552
resultpass passpass

Discussion. All three variants store every word gap as a real space character, so no extractor has to guess where words end. The totals only follow the amount of text. Neither the compressed style nor two columns changes how gaps are stored, since that is set by the typesetting engine and the template, not by the layout.

2 · Word-gap characters

What it tests. Test 1 checks how word gaps are stored in the PDF; this one checks what each engine returns for them. It lists any gap that comes back as something other than an ordinary space, above all the no-break space, which Java-based parsers do not treat as a separator and so read the two words as one. It only reports and never fails.

1c2p1c1p 2c1p
engines with only U+0020 gaps13/13 13/1313/13
gaps that would join words in Java000
resultreport clean report cleanreport clean

Discussion. All thirteen engines, Tesseract included, return a plain space for every word gap in all three variants, consistent with test 1: a real space character leaves nothing an engine could turn into a no-break space. No Java-based parser would run two words together.

3 · Icons extract as whitespace

What it tests. The contact line marks each field with a small icon (an envelope, the LinkedIn logo). Every icon must extract as a blank space, not as a stray letter or icon name that would corrupt the email or link next to it.

Icon font1c2p 1c1p2c1p
FontAwesome5Brands-Regular2 ok2 ok2 ok
FontAwesome5Free-Solid4 ok4 ok4 ok
resultpass passpass

Discussion. All three variants share the same contact block (two brand icons for LinkedIn and GitHub, four solid icons for the other fields), and all six icons extract as a plain space in every build. This is expected: icon extraction is set by the template, so layout choices cannot change it.

4 · PDF/A-2B conformance

What it tests. PDF/A is the archival PDF standard: every font embedded, no outside dependencies, a file that opens the same everywhere. The reference validator veraPDF confirms that each file really conforms to PDF/A-2B, since the PDF/A label in the metadata proves nothing on its own.

Check1c2p1c1p 2c1p
identificationPDF/A-2BPDF/A-2BPDF/A-2B
fonts embeddedallallall
output intentpresentpresentpresent
embedded files / forbidden actions / encryptionnonenonenone
PDF1.7, tagged1.7, tagged1.7, tagged
veraPDF 1.30.2pass passpass

Discussion. All three files are valid PDF/A-2B. The photo, about 2.6 MB of 1c2p and 2c1p, does not affect conformance, but it is the whole size difference, so the photo-less 1c1p suits portals that limit upload size.

5 · Text extraction

What it tests. The core test: every phrase must appear, complete, in each engine’s text. Differences that do not change the meaning (dash and quote styles, ligatures, letter case) are ignored.

Engine1c2p 1c1p2c1p
poppler91/9176/7673/73
poppler-raw91/9176/7673/73
poppler-layout91/9176/7649/73 advisory
mupdf91/9176/7673/73
pdfium91/9176/7673/73
pdfminer91/9176/7673/73
ghostscript91/9176/7650/73 advisory
pdfjs-old91/9176/7673/73
pdfjs-new91/9176/7673/73
pdfbox91/9176/7673/73
pdfbox-sorted91/9176/7649/73 advisory
tika91/9176/7673/73
tesseract (OCR)78/91 advisory70/76 advisory65/73 advisory

Two-column splicing in 2c1p

The three engines that lose phrases in 2c1p rebuild each printed line across the full page width, so a phrase that wraps in one column gets the other column’s text inserted where it breaks. Accordingly, 21–23 of the missing phrases are long, wrapping ones, spread over every section.

What poppler-layout extracts from 2c1p, verbatim
• Built a federated learning anomaly-detection and classification                Personal: Data-driven, Independent,  PoC for payment fraud detection, uncovering undetected fraud                   Organized, Systematic, Adaptable  and cutting detection time from weeks to minutes.

An Experience bullet (left column) and the Personal skills line (right column) merged line by line; ghostscript and pdfbox-sorted do the same.

Engine (2c1p)Lost By section
poppler-layout24experience 7, skills 4, education 3, volunteering 3, projects 2, courses 2, publications 2, about 1
ghostscript23experience 5, education 4, skills 4, volunteering 3, projects 2, courses 2, publications 2, about 1
pdfbox-sorted24experience 7, skills 4, education 3, volunteering 3, projects 2, courses 2, publications 2, about 1

Discussion: two columns. Every strict engine returns every phrase in all three variants, so the text stored in the PDF is complete whatever the layout. The only difference comes from two columns: in 2c1p, poppler-layout, ghostscript and pdfbox-sorted recover only 67–68% of the phrases. This is not a fault in the PDF: these engines cannot separate side-by-side columns, and they return 100% of both one-column variants. The compressed style causes no losses.

Tesseract: OCR misses by cause

Cause1c2p 1c1p2c1p
capital AI read as Al625
diacritics dropped (Chudá, Kučera)432
Scholar URL misread (9_uI59…)111
stray character (model → modell, Goldschmidt → Goldschmiadt)200
total13/916/768/73

Discussion: Tesseract OCR. Tesseract reads a rendered image of the page, not the text stored in the PDF, so its misses are OCR character errors rather than layout problems. As the table shows, they concentrate on the capital AI, which the sans-serif face makes hard to tell from Al, and on Czech and Slovak diacritics. Their number follows how often these strings occur, so the shortest variant, 1c1p, has the fewest. One stray character is new since the previous report: the 1c2p surname, set in Source Sans 3 instead of Source Sans Pro, comes back as Goldschmiadt. Rendering at 600 dpi does not help: the AI, diacritic and URL errors recur unchanged, and a few different stray characters replace the old ones, for 13 / 9 / 9 misses against 13 / 6 / 8 at 300 dpi.

Possible fix. The AI errors, and the Scholar URL, whose capital I comes back as l, have one cause: in Inter, the capital I and the lowercase l are both a plain vertical stroke, so OCR has to guess. The fix is a body typeface whose capital I and lowercase l differ in their regular shapes, for example an I with serifs or an l with a tail. The dropped diacritics are a different matter: they come from OCR running with an English-only language model, as in this test setup, and no change of font would fix them.

Practical relevance. In practice, the OCR errors matter little. ATS parsers work primarily with the text stored in the PDF, and since this file is tagged, PDF/A-2B-compliant and its text reads back completely, a fallback to pure OCR is unlikely. The OCR results are therefore for informational purposes only.

6 · Reading order

What it tests. Finding every phrase is not enough if they come back scrambled: an ATS that reads a date after the wrong employer files it under the wrong job. The test scores how much of each engine’s output keeps the order the text is stored in, from 0 to 1 (perfect). The body must score 1.00 in every strict engine; the contact header is only reported, since the order of contact fields carries no meaning.

Engine — body order score1c2p 1c1p2c1p
poppler1.00 (83/83)1.00 (68/68)0.86 (56/65) advisory
poppler-raw1.00 (83/83)1.00 (68/68)1.00 (65/65)
poppler-layout1.00 (83/83)1.00 (68/68)0.71 (29/41) advisory
mupdf1.00 (83/83)1.00 (68/68)1.00 (65/65)
pdfium1.00 (83/83)1.00 (68/68)1.00 (65/65)
pdfminer1.00 (83/83)1.00 (68/68)1.00 (65/65) advisory
ghostscript1.00 (83/83)1.00 (68/68)0.71 (30/42) advisory
pdfjs-old1.00 (83/83)1.00 (68/68)1.00 (65/65)
pdfjs-new1.00 (83/83)1.00 (68/68)1.00 (65/65)
pdfbox1.00 (83/83)1.00 (68/68)1.00 (65/65)
pdfbox-sorted1.00 (83/83)1.00 (68/68)0.71 (29/41) advisory
tika1.00 (83/83)1.00 (68/68)1.00 (65/65)
tesseract (OCR)1.00 (73/73)1.00 (64/64)1.00 (59/59)
header, all engines1.001.00 except poppler 0.88 (7/8) report1.00

In brackets: phrases in order / phrases found, so an engine that lost phrases is scored on fewer of them. The reference order comes from poppler-raw, which outputs the text exactly as stored.

poppler in 1c1p: LinkedIn moves to the end. poppler moves linkedin.com/in/… from the end of the first contact line to after the Scholar link on the second. The field is complete, only later in the output.
Reference order (poppler-raw): section headings in 2c1p
AboutExperienceEducationProjectsSkillsCourses & CertificationsSelected PublicationsVolunteering

The left column first, then the right one, as the PDF stores them.

What poppler extracts from 2c1p, verbatim (abridged)
About Skills Researcher with 9 years of applied research and R&D experience,… (rest of the About paragraph) Professional: Artificial Intelligence &…Tools: Python for AI (numpy, pandas,…Personal: Data-driven, Independent,…Languages: English (full professional…Experience

The right column’s Skills heading and block (highlighted) come out next to About, level with it at the top of the page, so an ATS would see Skills before Experience. These are all nine of poppler’s out-of-order phrases; every other section keeps its order.

Discussion. In both one-column variants every engine returns the body in document order, so order is a problem only in two columns. There, the seven engines that output text as stored (poppler-raw, mupdf, pdfium, both pdf.js builds, pdfbox, tika) keep a perfect 1.00, because the template stores each column as one continuous block. The engines that re-sort text by position depart from it to different degrees: pdfminer still reaches 1.00, poppler moves one block (0.86), and the three side-by-side engines also scramble what they recover (0.71).

Summary

TestPasses in all three? Where variants differ
real word spacesyesno difference: 770 / 557 / 552 real space characters
word-gap charactersreportno difference: U+0020 only, 13 of 13 engines
icons as whitespaceyesno difference: 6 of 6 icons in each
PDF/A-2Byesno difference
text extractionyes2c1p: 3 side-by-side engines lose 23–24 phrases (advisory); Tesseract misses 13 / 6 / 8 everywhere (advisory)
reading orderyes2c1p: 4 position-sorting engines at 0.71–0.86 (advisory); 1c1p: poppler moves LinkedIn in the header (reported)

Limits

Appendix

Environment

lualatexLuaHBTeX 1.24.0 (TeX Live 2026)
python3.14.6 (env/python)
pymupdf1.28.2
pypdfium25.13.0
pdfminer.six20260107
popplerpdftotext 26.01.0
ghostscript10.06.0
tesseract5.5.3
javaOpenJDK 25.0.3
pdfbox / tika3.0.5 / 3.2.3
nodev20.18.1
pdf.js2.5.207 (old), 4.10.38 (new)
veraPDF1.30.2

Reproduce

From tests/, one command per variant:

1c2p./pipeline.sh ../realizations/generic_research-1c2p/main.tex
1c1p./pipeline.sh ../realizations/generic_research-1c1p/main.tex
2c1p./pipeline.sh ../realizations/generic_research-2c1p/main.tex

Each run generates phrases.txt, builds <label>.pdf, runs run.sh and writes run.log and pipeline.json to tests/out/generic_research-<variant>/. The 600 dpi pass renders each built PDF with pdftoppm -r 600 -gray -png, runs tesseract on every page, and scores the output with the same check.py and order.py.

Changes to the tests since earlier reports. Phrase counts here are larger than in earlier reports and not comparable with them, for two reasons. Phrases are no longer deduplicated: a repeated phrase, or one inside a longer phrase, is now checked at each of its places. And text before the first section heading, which used to be skipped and left the About paragraph of 1c1p untested, is now checked too.

All three pipeline runs exited 0. Every figure here comes from their run.log, pipeline.json, order.json and the 13 per-engine extractions. The 2c1p loss breakdown and the Tesseract diagnosis were computed from those files with the tests’ own normalization.