Data and ML Stack Tools by Category

CategoryTools / TechnologiesPrimary Use
Machine LearningMLflow, Kubeflow, Feast, Ray, TensorFlow, PyTorch, Scikit-learn, XGBoost, LightGBMTraining, tracking, serving, feature store and MLOps
Data Science / AnalyticsPython, Jupyter, Pandas, Polars, NumPy, SciPy, R, Apache ArrowData exploration, statistical analysis and data preparation
BI / AnalyticsApache Superset, Metabase, Grafana, Power BI, TableauDashboards, data exploration and visualization
Data EngineeringApache Spark, Apache Flink, Apache Beam, Apache Kafka, dbtData processing, pipelines and transformations
Data Ingestion / CDCApache Kafka, Debezium, Kafka Connect, Apache NiFi, Airbyte, FivetranData ingestion, CDC and data integration
Workflow OrchestrationApache Airflow, Dagster, Prefect, Argo WorkflowsPipeline and workflow orchestration
Data Transformationdbt, Apache Spark, SQLMesh, DataformData transformation, modeling and ELT
Data QualityGreat Expectations, Soda, dbt Tests, DeequData validation and quality checks
Data Catalog / GovernanceOpenMetadata, DataHub, Apache Atlas, Amundsen, OpenLineageData catalog, lineage and governance
Data Lake StorageAmazon S3, MinIO, Google Cloud Storage, Azure Data Lake StorageObject storage and data lake
File FormatsApache Parquet, Apache ORC, Apache Avro, Apache ArrowData storage and interchange formats
Open Table FormatsApache Iceberg, Apache Hudi, Delta Lake, Apache PaimonACID tables, schema evolution, snapshots and time travel
Lakehouse CatalogApache Polaris, Project Nessie, AWS Glue Catalog, Hive MetastoreTable catalog and metadata management
Query EnginesTrino, Presto, DuckDB, ClickHouse, Apache Doris, StarRocksSQL analytics and OLAP
Lakehouse PlatformsDatabricks, Dremio, Starburst, SnowflakeData lakehouse and analytics platforms
Feature StoreFeast, Hopsworks, TectonML feature management and serving
ML PipelinesKubeflow Pipelines, MLflow, Airflow, Metaflow, FlyteMachine learning pipelines
Model ServingKServe, Seldon, NVIDIA Triton, Ray Serve, BentoML, vLLMModel inference and serving
Experiment TrackingMLflow, Weights & Biases, Neptune, CometExperiment tracking and metrics
Model RegistryMLflow, Kubeflow, Vertex AI, SageMaker, Azure MLModel versioning and lifecycle management
Vector DatabasesQdrant, Milvus, Weaviate, pgvector, PineconeEmbeddings, semantic search and RAG
Streaming / Event ProcessingApache Kafka, Apache Flink, Redpanda, Apache PulsarReal-time data processing
Data WarehouseSnowflake, BigQuery, Redshift, ClickHouse, DuckDBStructured analytics and OLAP
Data ObservabilityOpenTelemetry, Grafana, Prometheus, OpenLineage, MarquezData and pipeline observability

Reference architectures #

Three ways to wire the tools above into a working stack. All of them have the same three layers, ingestion, processing and storage, plus whatever consumes the result.

1. Batch lakehouse #

Nightly and hourly ELT. The default when latency is measured in hours.

flowchart TB
  subgraph ing["Ingestion"]
    airbyte["Airbyte"]
    debezium["Debezium + Kafka Connect"]
  end
  subgraph proc["Processing"]
    spark["Apache Spark"]
    dbt["dbt"]
    ge["Great Expectations"]
  end
  subgraph sto["Storage"]
    iceberg["Apache Iceberg"]
    s3["MinIO / S3"]
    polaris["Apache Polaris"]
  end
  subgraph out["Consumption"]
    trino["Trino"]
    superset["Apache Superset"]
  end

  airbyte --> spark
  debezium --> spark
  spark --> dbt --> ge --> iceberg
  s3 -.- iceberg
  polaris -.- iceberg
  iceberg --> trino --> superset

Apache Airflow sits outside the picture, triggering every arrow above.

2. Streaming #

Seconds instead of hours. Two branches out of one job: the table for history, the OLAP store for the dashboard.

flowchart TB
  subgraph ing["Ingestion"]
    debezium["Debezium"]
    kafka["Apache Kafka"]
  end
  subgraph proc["Processing"]
    flink["Apache Flink"]
  end
  subgraph sto["Storage"]
    paimon["Apache Paimon on S3"]
    clickhouse["ClickHouse"]
  end
  subgraph out["Consumption"]
    trino["Trino"]
    grafana["Grafana"]
  end

  debezium --> kafka --> flink
  flink --> paimon --> trino
  flink --> clickhouse --> grafana

Airflow is still needed here, not to schedule the job but to run the compaction and snapshot expiration the table needs because the job never stops writing.

3. ML and inference #

Features written by the same pipeline that feeds training, and a model served out of a registry.

flowchart TB
  subgraph ing["Ingestion"]
    kafka["Apache Kafka"]
  end
  subgraph proc["Processing"]
    flink["Apache Flink"]
    spark["Apache Spark"]
  end
  subgraph sto["Storage"]
    feast["Feast"]
    delta["Delta Lake on S3"]
    qdrant["Qdrant"]
  end
  subgraph train["Training"]
    ray["Ray"]
    mlflow["MLflow Registry"]
  end
  subgraph serve["Serving"]
    kserve["KServe"]
    vllm["vLLM"]
  end

  kafka --> flink
  kafka --> spark
  flink --> feast
  flink --> qdrant
  spark --> delta
  delta --> ray --> mlflow --> kserve
  feast --> kserve
  qdrant --> vllm

Official repositories #

Kafka Connect ships inside apache/kafka, dbt Tests inside dbt-core and Ray Serve inside ray-project/ray, so they share the repositories above.

No public repository, these are proprietary or managed services: Power BI, Tableau, Fivetran, Amazon S3, Google Cloud Storage, Azure Data Lake Storage, AWS Glue Catalog, Databricks, Starburst, Snowflake, BigQuery, Redshift, Tecton, Pinecone, Weights & Biases, Neptune, Comet, Vertex AI, SageMaker and Azure ML.

"only knowledge frees man" — E.C.