Drop in messy Excel, CSV or PDF files. Get back a normalized roster of names, IDs, dates and contact data — with a confidence score on every row.
| DEMO CO — SAMPLE DATA | |||
| Payroll export (fictional) | |||
| EMP_NO | EMPL_NM | SOC_SEC | BRTH_DT |
| 1001 | DOE, JANE B | 000-12-3456 | 01/14/90 |
| 1002 | SMITH, ALEX J | 000-23-4567 | 07/22/85 |
| 1003 | GARCIA, MARIA | 000-34-5678 | 03/09/92 |
Employee Identification Number
1001
Full Name (Last, First M.I.)
DOE, JANE B.
Social Security Number
000-12-3456
Full Date of Birth (MM/DD/YYYY)
01/14/1990
| ACME Corp | |||
| Q4 Report | |||
| EMP_ID | NAME | SSN | |
| 1001 | John S. | 123-XX | john@ |
| employee_id | full_name | ssn | |
| 1001 | John Smith | 123-45-6789 | john@acme.com |
| 1002 | Jane Doe | 234-56-7890 | jane@acme.com |
| 1003 | Bob Johnson | 345-67-8901 | bob@acme.com |
Real-world files, however messy. These are the shapes CyberInci handles out of the box — every preview below is fictional sample data.
| GLOBEX INC — INTERNAL USE | ||
| EMP_NO | EMPL_NM | SOC_SEC |
| 1001 | DOE, JANE B | 000-12-3456 |
| 1002 | SMITH, ALEX J | 000-23-4567 |
Header found on row 3 — branding skipped
Messy Excel exports
Branding rows, blank rows and logos above the real table — the header is found automatically, wherever it hides.
Download this sample3 sheets detected — all extracted
Multi-sheet workbooks
Every sheet is processed, scored and combined — no copy-pasting tabs together first.
Employee Number,Name of Employee,Social Sec #,Date Birth
1001,DOE JANE B,000-12-3456,01/14/1990
CSVs with human headers
Wordy or abbreviated headers are matched to the canonical schema by the fuzzy and AI tiers — no renaming needed.
Download this sample| NAME | ID | DOB |
| DOE, J | 1001 | 01/14/90 |
| SMITH, A | 1002 | 07/22/85 |
1 table found on page 2
PDFs with tables
Tables inside PDF reports are detected and extracted directly — same pipeline, same confidence scores.
Needs the free Tesseract OCR add-on
Scanned PDFs (OCR)
Image-only scans are read with the optional OCR add-on, then extracted like any other document.
| Employee ID? | Name? | SSN? |
| 1001 | DOE, JANE B | 000-12-3456 |
| 1002 | SMITH, ALEX J | 000-23-4567 |
No header row — columns inferred from values
Files with no headers at all
When there is no header row, columns are inferred from the data itself — IDs, names, SSNs and dates are recognized by shape.
Employee rosters, payroll reports, donor lists. Whatever the source, we normalize it.
Finds the real header row in branded, multi-row, or headerless files. Works with messy exports from any system.
Exact cache + fuzzy + sentence embeddings + LLM. Persistent learning means faster processing over time.
Confidence score on every row, every file. Routes outputs to Accept, Review, or Reject automatically.
Three steps from messy spreadsheet to normalized, validated PI data.
Drop your Excel or CSV files. Any format, any structure, any mess.
5-tier pipeline extracts PI columns with confidence scoring.
Get normalized output instantly, or review flagged files.