Tag
This article explains how TimescaleDB's hypercore engine achieves up to 98% compression for time-series data using columnar storage and specialized algorithms like delta encoding and Gorilla XOR, and contrasts it with PostgreSQL's TOAST.
GreptimeDB v1.1.0 is released, offering up to 97% faster PromQL queries, 20-40% lower overall query times, and up to 4.5x improvement on TSBS scan-heavy queries, along with online repartitioning for existing tables.
LangChain announces SmithDB, a purpose-built distributed database for agent observability that powers LangSmith, offering up to 12x performance improvements and support for complex agent trace queries.
This tweet introduces various development features provided by Cloudflare, including object storage R2, backend API Workers, AI gateway AI Gateway, containers, cache KV, database D1, and PostgreSQL connection HyperDrive, emphasizing their low cost, rich features, and generous free tier.
This post compares write-heavy sysbench performance of modern PostgreSQL (versions 15-19) and MySQL 8.4 on a large server, finding that InnoDB generally outperforms PostgreSQL on write throughput and shows less variation.
China open-sourced Zvec, an in-process vector database that runs inside apps without servers, supporting billions of vector searches in milliseconds and battle-tested at Alibaba scale.
A free, open-source tool that converts SQL CREATE TABLE statements into interactive entity-relationship diagrams locally in the browser, supporting multiple SQL dialects.
This research explores methods to determine the source table and column for each result column in arbitrary SQLite queries, using SQLite's internal column metadata API accessed via Python's apsw library or a ctypes bridge, with applications for tools like Datasette.
A blog post introducing Emacs rec mode as a plain-text database system, used for tracking books and integrating with Org mode.
This paper describes Scuba, a distributed in-memory database system developed at Facebook for real-time analytics and data exploration.
Oracle's AI Database now includes vector store functionality for embeddings-based image search, showcasing innovative features that make it a unified data storage solution.
Soma-SQL proposes an autonomous method to resolve multi-source ambiguity in natural language to SQL translation using synthetic query logs and ambiguity-driven execution probing, achieving 13% improvement in execution accuracy over state-of-the-art baselines.
A hands-on introduction to PostgreSQL using annotated SQL examples, covering basics to advanced topics.
PgDog, an open-source proxy that makes Postgres horizontally scalable, has raised $5.5M in funding from Basis Set, YC, and others. The tool is already serving over 2M queries per second in production.
ICE denies having a protester database, but a letter to Congress reveals the agency collects information on individuals involved in protests, including those not arrested, amid concerns over surveillance of U.S. citizens.
There is a meticulously compiled data engineering interview question bank on GitHub called data-engineering-interview-questions, containing over 2,000 questions covering databases, big data frameworks, cloud platforms, data visualization, and other core areas.
PostgreSQL documentation introduces property graphs, a SQL/PGQ feature that allows querying relational data using graph pattern matching syntax, defined as read-only views over tables.
Introduces UniQL, a human-verified executable benchmark for cross-dialect text-to-SQL evaluation, addressing the lack of dialect diversity in existing benchmarks like Spider and BIRD.
Erxi is welcomed as a new committer to GreptimeDB after shipping several improvements including meta KV write guards, flush-reason propagation, and CSV COPY fixes.
This article provides a technical breakdown of how the project management tool Linear achieves its fast performance by using a browser-side database (IndexedDB), local-first mutations, and a sync engine, eliminating network latency from user interactions.