Any document. Clean, compliance-ready data.

Drop in messy Excel, CSV or PDF files. Get back a normalized roster of names, IDs, dates and contact data — with a confidence score on every row.

sample_payroll_export.xlsxdemo · sample data
DEMO CO — SAMPLE DATA   
Payroll export (fictional)   
    
EMP_NOEMPL_NMSOC_SECBRTH_DT
1001DOE, JANE B000-12-345601/14/90
1002SMITH, ALEX J000-23-456707/22/85
1003GARCIA, MARIA000-34-567803/09/92
Branding rows detected & skipped · header row 4 · score 0.97
5-tier AI
Canonical outputACCEPT · 1.00

Employee Identification Number

1001

cache

Full Name (Last, First M.I.)

DOE, JANE B.

cache

Social Security Number

000-12-3456

fuzzy

Full Date of Birth (MM/DD/YYYY)

01/14/1990

embedding
messy_data.xlsx
 ACME Corp  
 Q4 Report  
EMP_IDNAMESSNEMAIL
1001John S.123-XXjohn@
Cache
Fuzzy
Embedding
LLM
Infer
normalized_pi.xlsx
employee_idfull_namessnemail
1001John Smith123-45-6789john@acme.com
1002Jane Doe234-56-7890jane@acme.com
1003Bob Johnson345-67-8901bob@acme.com
Confidence: 0.9280/81 rows keptACCEPT

What can it extract?

Real-world files, however messy. These are the shapes CyberInci handles out of the box — every preview below is fictional sample data.

quarterly_roster.xlsxxlsx
GLOBEX INC — INTERNAL USE
EMP_NOEMPL_NMSOC_SEC
1001DOE, JANE B000-12-3456
1002SMITH, ALEX J000-23-4567

Header found on row 3 — branding skipped

Messy Excel exports

Branding rows, blank rows and logos above the real table — the header is found automatically, wherever it hides.

Download this sample
regional_offices.xlsxls
EastWestCentral

3 sheets detected — all extracted

Multi-sheet workbooks

Every sheet is processed, scored and combined — no copy-pasting tabs together first.

contact_list.csvcsv

Employee Number,Name of Employee,Social Sec #,Date Birth

1001,DOE JANE B,000-12-3456,01/14/1990

"Social Sec #" → SSN"Date Birth" → DOB

CSVs with human headers

Wordy or abbreviated headers are matched to the canonical schema by the fuzzy and AI tiers — no renaming needed.

Download this sample
benefits_report.pdfpdf
NAMEIDDOB
DOE, J100101/14/90
SMITH, A100207/22/85

1 table found on page 2

PDFs with tables

Tables inside PDF reports are detected and extracted directly — same pipeline, same confidence scores.

old_records_scan.pdfpdf
OCR

Needs the free Tesseract OCR add-on

Scanned PDFs (OCR)

Image-only scans are read with the optional OCR add-on, then extracted like any other document.

export_no_headers.csvcsv
Employee ID?Name?SSN?
1001DOE, JANE B000-12-3456
1002SMITH, ALEX J000-23-4567

No header row — columns inferred from values

Files with no headers at all

When there is no header row, columns are inferred from the data itself — IDs, names, SSNs and dates are recognized by shape.

Built for messy data

Employee rosters, payroll reports, donor lists. Whatever the source, we normalize it.

Smart header detection

Finds the real header row in branded, multi-row, or headerless files. Works with messy exports from any system.

Layered column mapping

Exact cache + fuzzy + sentence embeddings + LLM. Persistent learning means faster processing over time.

Per-row plausibility

Confidence score on every row, every file. Routes outputs to Accept, Review, or Reject automatically.

How it works

Three steps from messy spreadsheet to normalized, validated PI data.

01

Upload

Drop your Excel or CSV files. Any format, any structure, any mess.

02

Process

5-tier pipeline extracts PI columns with confidence scoring.

03

Download / Review

Get normalized output instantly, or review flagged files.

Ready to normalize your data?

Upload your first file and see the 5-tier pipeline in action.