Skip to main content

Initiative #152

← AI initiative registry

Data Validation Tool

Listed in the registry Pilot Original language : English

About the initiative

Description of the initiative

Use large language models to help detect errors in data transcribed using optical character recognition. Apply validation rules to propose corrections and normalize data into a standardized format.

AI family

Automation and Decision Support

Bucket rationale Not translated

Automatisation de la validation et de la correction de données structurées.

Lifecycle

  1. Ideation

    Stage status : Completed

    June 1, 2025

  2. Assessment

    Stage status : Completed

  3. Approved

    Stage status : Completed

  4. Development

    Stage status : Completed

  5. Pilot

    Stage status : In progress

  6. Production

    Stage status : Upcoming

  7. Archived

    Stage status : Upcoming

Schedule for this initiative

  • Completed
  • In progress
  • Planned
  • Not applicable
  • Today
2025 Jun Jul Aug Sep Oct Nov Dec 2026 Jan Feb Mar Apr May Jun Jul Aug Sep
  1. Ideation

A similar need?

Nobody has come forward yet. If your team faces the same problem, say so: it helps bring initiatives together and share a solution.

Your email is never published: only the AI Enablement team sees it, to connect you.

Contacts

Requester
Smith, Melannie
Sector contact
Smith, Melannie
Subject matter expert
Winegardner, Amanda
Smith, Melannie

Data and tools

Sector
Ecosystems and Oceans Science
Region
National Capital Region
Could this initiative be shared with the Treasury Board Secretariat?
Yes
Submitted on
June 1, 2025

Primary users

  • Departmental employees

Government priorities and declaration

Declared to TBS Recorded as transmitted to TBS; not editable here.

Declared name of the AI system
Data Validation Tool (See B6, B22)
TBS registry identifier
2526-DFO-MPO-013
Status declared to TBS
In development
Sent to TBS on
October 10, 2025
Purpose of the system, as declared
Problem: Digitizing handwritten or non-machine-readable documents using optical character recognition (OCR) often introduces transcription errors, which can compromise data quality and usability. These inaccuracies make it difficult to rely on the extracted information for analysis or integration into standardized systems. Objective for AI: Leverage large language models to detect and correct OCR-induced errors by applying intelligent validation rules. The goal is to propose accurate corrections and normalize the extracted data into a consistent, standardized format, ensuring higher reliability and interoperability across systems.