Document ingestion
Processed PDFs, Excel, scanned documents and images using OCR with Google Document AI and Gemini for document classification and structured data extraction. Heterogeneous KID layouts, multi-language tables and noisy scans all funnelled into a single normalised schema.
Validation & enrichment
Delivered asset categorisation, ISIN matching, web enrichment, KID-document analysis and cost extraction across heterogeneous sources. Each extracted field is paired with a confidence score and a one-click fix path for the analyst.
Database & GDPR
Contributed to the database design for asset storage and historical tracking, ensuring data traceability and GDPR-compliant processing throughout the pipeline.
Reporting & API
Generated professional PDF and Excel benchmark reports with "Old vs New" validation; built a prototype application and an API access layer for the full ingestion-to-reporting workflow.