DATA & ANALYTICS

Data science, analytics, databases, and visualization. The infrastructure layer that every data-driven product is built on.

50 entriesranked by early signal score

§ 1 — catalog entries

MatrixOne is a single database that handles storing, searching, and analyzing data all in one place, including the ability to search by meaning (like how AI understands…

why it matters: Builders creating AI-powered products typically need to stitch together multiple separate databases and tools, which adds cost and complexity …

2.0k★331⑂135 contributorsGo

scipy/scipy

54/100

Hot

SciPy is a free, open-source software library that gives Python programmers a ready-made toolkit for solving complex mathematical and scientific problems — things like…

why it matters: With nearly 15,000 stars and close to 1,900 contributors, SciPy is essentially the standard plumbing beneath countless data science, research, and…

15.1k★6.0k⑂1.9k contributorsPython

pgGraph lets you run powerful relationship and network queries — the kind normally requiring a specialized graph database — directly on top of your existing PostgreSQL…

why it matters: Builders typically face an expensive, risky choice between sticking with a familiar database or adopting a whole new graph database system just to…

1.1k★86⑂3 contributorsRust

numpy/numpy

53/100

Hot

NumPy is a foundational Python library that makes it fast and easy to work with large collections of numbers and data — think spreadsheets on steroids that computers can…

why it matters: Nearly every AI, data science, and analytics product built in Python depends on NumPy, making it one of the most critical pieces of shared…

32.9k★12.9k⑂2.1k contributorsPython

ClickHouse is an open-source database built specifically for analyzing massive amounts of data at lightning speed, returning results in real-time rather than making you…

why it matters: As user expectations shift toward real-time everything, products that can surface instant insights from data have a significant competitive edge over…

50.2k★9.0k⑂3.2k contributorsC++

afni/afni

52/100

Hot

AFNI is a comprehensive software toolkit used by neuroscientists to process, analyze, and visualize brain scan images, including the functional MRI scans (brain imaging…

why it matters: Brain imaging research underpins a massive and growing market spanning clinical neurology, mental health diagnostics, and neurotechnology, and AFNI is…

196★122⑂81 contributorsC

Foxglove SDK is a toolkit that lets robotics and engineering teams record, stream, and visually explore complex sensor data — think camera feeds, GPS tracks, and sensor…

why it matters: As robotics, autonomous vehicles, and industrial automation become major investment areas, teams need better tools to understand and debug what their…

311★110⑂45 contributorsRust

apache/spark

51/100

Hot

Apache Spark is a powerful open-source platform that lets companies process and analyze massive amounts of data extremely fast — think analyzing billions of records in…

why it matters: With over 44,000 stars and 3,400 contributors, Spark is effectively the industry standard for large-scale data processing, meaning any data-heavy…

44.1k★29.4k⑂3.4k contributorsScala

simdjson is a software library that reads and processes JSON data — the universal format used to send information between apps and servers — dramatically faster than…

why it matters: When your product handles large volumes of data, the speed at which you can read and process that data directly affects your infrastructure costs and…

24.3k★1.3k⑂199 contributorsC++

Apache Airflow is an open-source platform that lets teams build, schedule, and monitor automated workflows — think of it as a smart traffic controller for your data…

why it matters: For any company building data-driven products or AI features, Airflow is often the backbone that keeps everything running reliably — making it a…

47.0k★17.9k⑂4.7k contributorsPython

Elasticsearch is a powerful search and data engine that lets companies instantly search through massive amounts of data and find relevant results in near real-time …

why it matters: With nearly 78,000 stars and over 2,500 contributors, Elasticsearch is one of the most widely adopted search engines in the world, meaning builders…

78.2k★26.1k⑂2.6k contributorsJava

Indicator is a free, open-source toolkit for building automated stock and financial market analysis tools, offering over 80 pre-built formulas (like moving averages and…

why it matters: For fintech founders and quant trading teams, this library dramatically shortens the time to build a working trading or market analysis product — what…

1.8k★279⑂14 contributorsGo

PostHog is an all-in-one platform that helps product teams understand how people use their software — tracking user behavior, replaying real sessions, running…

why it matters: As AI agents become part of the product development process, having all your user data in one place means those agents can act on richer context …

40.0k★3.5k⑂444 contributorsPython

Pandas is a widely-used Python library that makes it easy to organize, clean, and analyze large sets of structured data — think of it like a supercharged spreadsheet that…

why it matters: With nearly 50,000 stars and over 4,000 contributors, pandas is effectively the industry standard for data manipulation, meaning any product that…

49.9k★20.5k⑂4.3k contributorsPython

Grafana is a free, open-source tool that pulls data from virtually any source and turns it into visual dashboards — charts, graphs, and alerts — so teams can monitor the…

why it matters: With 75,000+ stars and nearly 3,000 contributors, Grafana has become the de facto standard for business and product monitoring, meaning building on or…

77.0k★14.8k⑂3.0k contributorsTypeScript

cogent3 is a Python library that helps scientists analyze DNA and genomic sequence data, enabling researchers to study how species evolve and compare genetic information…

why it matters: Genomics and biological data analysis is a rapidly growing field powering drug discovery, personalized medicine, and agricultural biotech — tools like…

137★68⑂92 contributorsPython

PostgreSQL is one of the world's most popular open-source databases, used to store, organize, and retrieve data for virtually any type of application — from small…

why it matters: With 22,000+ stars and nearly 6,000 forks, PostgreSQL's massive adoption means it's a safe, battle-tested foundation for almost any product you're…

22.3k★5.9k⑂58 contributorsC

SimulationCraft is a powerful simulator for World of Warcraft that lets players model and predict how much damage their characters will deal under various combat…

why it matters: This project demonstrates the strong demand for data-driven decision-making tools within gaming communities, where players are willing to engage with…

1.6k★786⑂500 contributorsC++

apache/beam

46/100

Hot

Apache Beam is an open-source framework that lets developers write a single data processing program that can handle both large historical datasets and live, real-time…

why it matters: For companies building data-heavy products, Beam dramatically reduces the engineering cost of switching cloud providers or scaling up data pipelines…

8.7k★4.7k⑂2.0k contributorsJava

duckdb/duckdb

46/100

Hot

DuckDB is a fast, lightweight database that runs directly inside your application — no separate server required — and is built specifically for analyzing large amounts of…

why it matters: As data-driven products become the norm, DuckDB lets small teams run powerful data analysis without the cost and complexity of traditional data…

41.8k★3.8k⑂709 contributorsC++

OpenAnalytics is a free, open-source website analytics tool that tracks visitor behavior and revenue without using cookies or storing personal data, making it a…

why it matters: With growing privacy regulations and user distrust of invasive tracking, builders can deploy this instead of Google Analytics to stay compliant…

589★48⑂SoloTypeScript

DBeaver is a free desktop application that lets you connect to and manage virtually any database — from MySQL and PostgreSQL to Snowflake and Oracle — through a single…

why it matters: With over 51,000 GitHub stars and support for 100+ databases, DBeaver has become a default tool for data-heavy teams, meaning your developers and…

51.9k★4.4k⑂468 contributorsJava

Apache Iceberg Rust is an open-source project that helps companies manage and organize massive amounts of data stored in data lakes (large, centralized repositories where…

why it matters: As companies accumulate ever-growing volumes of data, the tools they use to manage it become a critical competitive advantage — faster, more reliable…

1.4k★582⑂191 contributorsRust

Vibe-Research is an open-source personal investment research dashboard that pulls together market data, financial reports, news, and portfolio tracking for Chinese…

why it matters: As AI-powered investing tools go mainstream, there's a growing market for self-hosted, privacy-conscious alternatives to expensive research platforms…

2.6k★527⑂SoloTypeScript

dbt-labs/dbt

45/100

Hot

dbt is a tool that helps data teams clean, organize, and transform raw business data into reliable, analysis-ready reports — using the same disciplined, review-based…

why it matters: dbt has become one of the most widely adopted tools in the modern data stack, with nearly 14,000 GitHub stars and a large community, meaning it sits…

13.9k★2.6k⑂452 contributorsRust

trinodb/trino

45/100

Hot

Trino is an open-source engine that lets companies ask complex questions across massive amounts of data stored in different places, all using standard SQL — the same…

why it matters: For any company sitting on large amounts of data spread across multiple storage systems, Trino eliminates the need to build expensive, time-consuming…

13.3k★3.8k⑂1.1k contributorsJava

Metabase is an open-source business intelligence tool that lets anyone in a company explore data, build charts, and create dashboards without needing to know how to write…

why it matters: With nearly 48,000 GitHub stars and half a million users, Metabase has become the default self-hosted analytics layer for startups that want to avoid…

49.5k★6.9k⑂499 contributorsClojure

apache/kafka

44/100

Hot

Apache Kafka is a high-speed messaging system that lets companies move massive amounts of data between different applications and services in real time — think of it as a…

why it matters: With 33,000+ stars and adoption across major enterprises, Kafka has become the de facto backbone for real-time data movement, meaning any product that…

33.9k★15.5k⑂1.8k contributorsJava

OpenSearch is a free, open-source search and analytics engine that lets you add powerful search functionality to your products — think searching through massive amounts…

why it matters: With over 13,000 stars and 2,000+ contributors, OpenSearch has become a serious community-backed alternative to expensive enterprise search tools…

13.8k★3.0k⑂2.2k contributorsJava

apache/flink

44/100

Hot

Apache Flink is a powerful data processing engine that can handle massive streams of information in real time — think processing millions of events per second as they…

why it matters: For any product that competes on speed and freshness of data — real-time personalization, live analytics, instant alerts — Flink is the kind of…

26.4k★14.0k⑂2.1k contributorsJava

This is a collection of ready-made pipeline templates from Google that automate common data tasks in the cloud — like moving data between storage systems, creating…

why it matters: For builders, this dramatically cuts the time and engineering effort needed to set up data workflows between cloud services, which is often a costly…

1.3k★1.1k⑂242 contributorsJava

PolicyEngine US is an open-source tool that models how US federal and state tax and benefit programs work, letting you calculate how policy changes would affect people's…

why it matters: Builders creating financial planning tools, benefits eligibility apps, or policy analysis platforms can plug into this instead of building complex…

165★214⑂145 contributorsPython

This project is a university course repository teaching students how to process and analyze massive datasets quickly using powerful computing systems, with a focus on…

why it matters: As data volumes explode across every industry, the ability to process large datasets fast is becoming a core competitive advantage — and this…

151★140⑂90 contributorsJupyter Notebook

Tracearr is a unified dashboard that lets you monitor all your personal media servers — Plex, Jellyfin, and Emby — in one place, showing who is watching what in real time…

why it matters: As self-hosted media servers grow in popularity among privacy-conscious users, tools that add oversight and control become essential — and account…

2.7k★118⑂19 contributorsTypeScript

mongodb/mongo

43/100

Hot

MongoDB is one of the world's most popular open-source databases, used to store and retrieve data for applications of all sizes — from startups to Fortune 500 companies…

why it matters: With nearly 29,000 stars and over 1,400 contributors, MongoDB's widespread adoption means a massive talent pool, extensive tooling, and a proven track…

28.6k★5.8k⑂1.4k contributorsC++

This project connects ClickHouse (a fast database designed for analyzing large amounts of data quickly) to Grafana (a popular tool for creating visual dashboards and…

why it matters: As more companies adopt ClickHouse for high-speed data analysis, this plugin reduces the time and cost of building internal analytics dashboards by…

223★135⑂83 contributorsTypeScript

Matplotlib is a Python library that lets developers turn raw data into charts, graphs, and visualizations — from simple line graphs to complex animated figures suitable…

why it matters: With nearly 1,900 contributors and over 23,000 stars, Matplotlib is effectively the industry standard for data visualization in Python, meaning any…

23.3k★8.5k⑂1.9k contributorsPython

Legend Studio is an open-source visual workspace for designing and managing data models — essentially a sophisticated diagram and editing tool that lets teams define how…

why it matters: For any company dealing with complex data — financial services, healthcare, enterprise software — having a standardized way to define and communicate…

111★150⑂81 contributorsTypeScript

PySPEDAS is an open-source Python toolkit that lets researchers download, analyze, and visualize data from over 30 space and Earth observation missions — covering…

why it matters: As commercial space, satellite communications, and space weather risk monitoring become serious industries, tools that make government space data…

204★78⑂50 contributorsPython

OpenElectricity is an open platform that collects and organizes Australia's public energy market data — think electricity generation, consumption, and grid activity — and…

why it matters: Anyone building energy monitoring tools, climate tech products, or investment dashboards for the Australian market can skip months of data wrangling…

131★37⑂12 contributorsPython

CourtListener is a free, searchable archive of U.S. legal data — including court opinions, judge records, financial disclosures, and federal case filings — that has been…

why it matters: Legal data is notoriously hard to access and expensive to license, making CourtListener a rare open dataset that legaltech startups, AI companies, and…

1.0k★274⑂125 contributorsPython

Velox is an open-source software library created by Meta that gives companies a high-performance engine for processing and querying large amounts of data, acting as the…

why it matters: Companies like Microsoft, ByteDance, and IBM are already using Velox as the engine inside their own data products, which means this is becoming a…

4.2k★1.6k⑂664 contributorsC++

This is the backend system that powers DefiLlama, a popular website that tracks and displays financial data across hundreds of decentralized finance (DeFi) platforms …

why it matters: DefiLlama is one of the most widely used data sources in the crypto industry, meaning this codebase underpins a tool that investors, founders, and…

227★1.4k⑂762 contributorsTypeScript

DefiLlama Adapters is a community-built collection of small code plugins that pull financial data from decentralized finance (DeFi) applications, allowing DefiLlama to…

why it matters: With over 6,900 forks and 405 contributors, this repository shows the scale of the DeFi ecosystem and how many teams actively want their projects…

1.3k★7.8k⑂5.6k contributorsJavaScript

Aptos Explorer is the official window into the Aptos blockchain, letting anyone look up transactions, account balances, and network activity in real time — similar to how…

why it matters: For any team building products on the Aptos blockchain, a reliable and open explorer is essential infrastructure — it's how users verify their…

124★151⑂51 contributorsTypeScript

Spellbook is a shared library of pre-built data queries that make it easier to analyze blockchain activity on Dune, a popular crypto data platform. Instead of each…

why it matters: With nearly 1,400 contributors, this project signals strong community demand for standardized crypto analytics, which is increasingly critical for…

1.5k★1.4k⑂713 contributorsPython

CODAP is a free, browser-based data analysis and visualization tool designed specifically for students and educators, allowing them to explore and make sense of data…

why it matters: Backed by NSF funding and integrated into multiple established educational programs, CODAP represents a growing market for accessible data literacy…

106★49⑂40 contributorsTypeScript

Delta Kernel RS is an open-source library that lets any data processing tool read from and write to Delta tables — a popular format for storing and managing large…

why it matters: As more companies bet on Delta Lake as their data storage standard, this library lowers the barrier for any team to build tools that integrate with…

363★218⑂75 contributorsRust

DefiLlama is the open-source codebase behind the leading analytics dashboard for decentralized finance, tracking how much money is flowing through over 6,000 financial…

why it matters: With 129 contributors and hundreds of forks, this is effectively the industry-standard data layer that DeFi products, investors, and journalists rely…

292★372⑂130 contributorsTypeScript

Apache Superset is a free, open-source platform that lets teams explore, analyze, and visualize data through interactive charts and dashboards — no coding required for…

why it matters: With nearly 75,000 stars and over 1,400 contributors, Superset has become the go-to open-source alternative to expensive business intelligence tools…

75.0k★18.4k⑂1.5k contributorsPython

§ 2 — other categories

form 27-b — subscription

THE TUESDAY BRIEFING

The repos that moved this week, why they matter, and what to watch next. One email. No noise.