#data-engineering (12)
- Data Quality Monitoring Tools: Complete Comparison 2026
Comprehensive comparison of leading data quality monitoring tools – features, performance, pros & cons, and when to use each for high-volume data and alert fatigue.
- Collaborative Data Management with Dolt Remotes and DoltHub
Learn how to enable collaborative data management with Dolt remotes and DoltHub, mastering data synchronization, sharing, and team workflows for versioned SQL databases.
- Meta's Petabyte-Scale Data Ingestion Migration: Technical Case Study
In-depth case study of Meta's massive data ingestion system migration – architecture, implementation, challenges, and results for enhancing reliability at scale.
- Building AI/ML Pipelines: From Data to Deployment
Explore the foundational concepts of AI/ML pipelines, from data ingestion and preparation to model training, deployment, and continuous monitoring, crucial for scalable AI applications.
- Building Robust Pipelines: From Ingestion to Vectorization
Explore the critical steps of data ingestion, preprocessing, and vectorization for multimodal AI systems, focusing on robust and high-performance pipeline design.
- AI-Powered Systems: Debugging Models & Data Pipelines
Master debugging techniques for AI models and data pipelines, covering data quality, model performance, prompt engineering, and observability in modern AI systems.
- 16. Project: Data Pipeline Testing with Python (Kafka & DB)
Build and test a simple data pipeline in Python using Testcontainers to spin up isolated Kafka and PostgreSQL instances. Learn end-to-end integration testing for complex data flows.
- Building Custom Connectors & Extensions
Learn how to extend MetaDatasetFlow with custom connectors and transformers for unique data management tasks.
- Monitoring & Observability for Data Pipelines
Learn how to monitor and observe data pipelines for high-quality, reliable data in machine learning projects.
- Integrate OpenZL into Data Pipelines Using SDDL
Integrate OpenZL into your data pipelines by defining data structures with SDDL, then efficiently compress and decompress data using practical Python examples.
- Integrating OpenZL with Existing Data Workflows
Learn how to integrate OpenZL into your data workflows for efficient storage and performance improvement.
- Ingesting & Harmonizing HS Code and Tariff Data
Learn how to build a robust data pipeline using Databricks Delta Live Tables to ingest, cleanse, and harmonize HS Code and tariff data into your Customs Trade Data Lakehouse.