Dmitriy Kononov.
Let’s talkContact

SOLUTIONS

Python for documents, data and repeatable processing

Data analysis, OCR, API pipelines, permitted web data collection, visualisation and automation.

Area of expertise

Business challenge → design → delivery

Concept illustration of the workflow

Make recurring data work repeatable

I use Python for analysis, automation, API integrations, visualisation, document extraction and web data collection. I aim for a process that can run again on the next batch of files, with inputs, transformations and exceptions available for review.

Define the input and the decision

We identify file formats, data sources, expected fields and the intended use of the output. An initial stage can process one representative dataset, validate it and export a structured result. Pandas supports tabular transformations; OCR and computer vision can help with scanned documents. The appropriate approach depends on actual samples and error tolerance.

For web data collection, source permissions and available APIs guide the method. Machine-learning or predictive analysis needs a suitable dataset and validation plan; we would establish those requirements before including it.

PDF extraction and campaign reporting

The PDF/OCR outline describes Tesseract, OpenCV, pdf2image, Pandas, batch processing and JSON/CSV/MariaDB outputs. The completed Mailwizz project joins API data, spreadsheets and registration logs.

Acceptance includes a reference sample, problematic records, reruns, logs and reconciliation of totals. If the output triggers a customer or financial action, review and correction boundaries must be explicit. For scheduled operation, infrastructure and recovery form part of the design.

The original project descriptions are in my earlier Python section. Describe the data and the result you need.