The Data Pipeline Failure Investigator

 Create a data pipeline investigator that reconstructs failures across scheduled jobs, ingestion processes, transformations, queues, and downstream reports. Accept pipeline definitions, execution logs, schema versions, timestamps, and data-quality results. Build a timeline showing where records stopped flowing, changed unexpectedly, duplicated, or arrived late. Trace downstream datasets and reports affected by the failure, identify plausible root causes with supporting evidence, and generate targeted validation queries and recovery steps. Distinguish confirmed causes from hypotheses and never recommend replaying data without checking duplication and idempotency risks.

Post a Comment

0 Comments