Hungarian Public Procurement Authority
Structuring five years of public procurement — and making it searchable in plain language.
A national authority sat on 100,000+ unstructured XML tender documents. We built the AI and data-mining pipeline that structures them, plus a RAG-based legal chatbot that answers case handlers with citations.
Context
Public procurement data decides where public money goes. The authority’s archive grew by thousands of XML filings a month, each structured differently, each legally meaningful.
The costly problem
Case handlers spent roughly 35 minutes per research query — opening filings one by one, cross-checking legal references manually, re-keying findings into reports.
Baseline
Measured over 4 weeks with 12 case handlers: 35 min median per query, 11% of answers later corrected after supervisor review, no coverage statistics at all.
Intervention
A hybrid pipeline: deterministic XML mapping where structure exists, ML models where it varies, and an LLM layer for legal-language retrieval. The chatbot answers in Hungarian with citations to the exact filing and paragraph, so every claim is inspectable.
Human responsibilities
- Case handlers approve every extracted field set before it enters the official register
- Legal officers review a weekly sample of chatbot answers; failures feed the evaluation set
- The authority’s IT team owns the deployment — we document and hand over, they operate
Measured outcome
Limitations
Savings are measured for research and retrieval work, not for legal judgment — that stays with the officers. The EU compliance scope covers the defined data-due-diligence fields, not a blanket certification.
Research queues cleared same-day instead of piling up; case handlers shifted time to substantive legal review. Training took two half-day sessions; adoption tracked weekly and held above 80% after month two.
Next step
Extending the same pipeline to award-decision documents and supplier-history analytics.