User-ranked databases, warehouses, and data infrastructure tools.
Airbyte review: open-source ELT platform with 600+ connectors, self-hosted or cloud. License terms, pricing, and Airbyte vs Fivetran/dbt for data pipelines.
Apache Airflow workflow orchestration review: DAGs, scheduling, monitoring. Best for data pipeline management.
Amazon Athena review: the serverless managed Presto/Trino service. $5 per TB scanned pricing, querying S3 with SQL, cost pitfalls, and Athena vs self-managed Trino.
BigQuery: Google Cloud, SQL queries, massive scale. Essential data infrastructure tool.
Apache Cassandra database review: distributed NoSQL, linear scalability, high availability. Built for massive scale.
ClickHouse review: open-source columnar OLAP database for real-time analytics on billions of rows. Pricing, use cases, and how it compares to Snowflake and DuckDB.
Azure Cosmos DB review: globally distributed, multi-model database. 99.999% availability SLA, turnkey global distribution.
Couchbase review: distributed NoSQL, key-value + document + search. High performance, mobile sync with Couchbase Lite.
Dagster review: open-source, asset-based data orchestrator built in Python. Apache 2.0 license, pricing, and Dagster vs Airflow for pipeline orchestration.
dbt: SQL-based transforms, version control, testing. Essential data infrastructure tool.
DuckDB review: the in-process 'SQLite for analytics'. Query Parquet and CSV files with SQL, free MIT license, MotherDuck cloud option, and DuckDB vs pandas compared.
AWS DynamoDB review: serverless NoSQL, single-digit millisecond latency, automatic scaling. AWS-native key-value store.
Elasticsearch review: distributed search engine, full-text search, log analytics. Real-time search at scale.
Fivetran: automated connectors, ELT platform, no-code integration. Essential data infrastructure tool.
Apache Flink review: stream processing framework. True real-time, stateful computations, exactly-once semantics.
Apache Hive review: SQL-on-Hadoop data warehouse. Query massive datasets with HiveQL. Legacy batch processing.
InfluxDB time-series database review: metrics, events, IoT data. Purpose-built for time-stamped data.
Apache Kafka review: distributed event streaming platform. Real-time data pipelines, event-driven architectures. Industry standard.
MongoDB document database explained: how the document model works, when documents beat rows and tables, pricing from free to Atlas, and honest limitations.
MySQL: popular open-source, web applications, LAMP stack. Essential data infrastructure tool.
Neo4j graph database review: relationship-first, Cypher query language. Built for connected data.
PostgreSQL database review: features, performance, extensions. Most advanced open-source relational database.
Presto SQL explained: how the distributed query engine works, Presto vs Trino, and managed Presto services (Athena, EMR, Starburst) with pricing compared.
Redis: caching, real-time applications, sub-millisecond latency. Essential data infrastructure tool.
Amazon Redshift: AWS native, columnar storage, parallel processing. Essential data infrastructure tool.
Snowflake review: how credit-based pricing really works, what a month actually costs, and a straight Snowflake vs Redshift vs BigQuery comparison table.
Apache Spark review: unified analytics engine for big data processing. Fast in-memory computation, batch and streaming.
Starburst review: managed Trino from the engine's creators. Galaxy SaaS pricing, Enterprise self-managed option, Gravity data products, and Starburst vs Athena.
TimescaleDB review: hypertables explained, real 90%+ compression numbers, TimescaleDB vs InfluxDB, and what the TigerData rebrand means for the extension.
Trino review: the actively developed Presto fork for distributed SQL. Connectors, pricing, Iceberg support, and how it compares to Athena and Starburst.