withmarbleapp/os-taxonomy

Marble has open-sourced a detailed map of everything children learn during their primary school years, broken into 1,590 individual concepts — things like 'building sentences' or 'apparent brightness of stars' — with a web of connections showing which ideas a child must understand before they can learn the next one. The dataset is freely available as structured files and covers eight subjects, aligned to major curriculum standards like Common Core and the UK National Curriculum.

4.2k749SoloJavaScriptsource ↗

§ 1 — what it does

Marble has open-sourced a detailed map of everything children learn during their primary school years, broken into 1,590 individual concepts — things like 'building sentences' or 'apparent brightness of stars' — with a web of connections showing which ideas a child must understand before they can learn the next one. The dataset is freely available as structured files and covers eight subjects, aligned to major curriculum standards like Common Core and the UK National Curriculum.

§ 2 — why it matters

Any team building an edtech product — tutoring apps, adaptive learning platforms, AI-powered homework tools — typically has to spend years and significant resources creating this kind of curriculum structure from scratch, so having it open-sourced removes a major barrier to entry. This is the kind of foundational dataset that could become the backbone of a new generation of personalized learning products, much like open maps data enabled a wave of location-based startups.

§ 4 — related entries

4 entries

numpy/numpy

61/100

Hot

NumPy is the foundational Python library for working with large collections of numbers and mathematical data, enabling everything from basic calculations to complex simulations at high speed. It acts as the backbone that almost every data science and AI tool in Python is built on top of, making it essential infrastructure for any software that processes numerical information.

why it matters: With over 32,000 stars and 2,100 contributors, NumPy is effectively a universal dependency in the AI and data ecosystem — if your product touches machine learning, data analysis, or scientific computing, it almost certainly relies on NumPy under the hood. Builders should understand that investing in or building on this ecosystem means standing on extremely stable, widely adopted infrastructure, but also that any major changes to NumPy can ripple across thousands of downstream products.

32.6k12.7k2.1k contributorsPython

Apache Airflow is an open-source platform that lets teams build, schedule, and monitor automated workflows — think of it as a smart traffic controller for your data pipelines, ensuring the right tasks run in the right order at the right time. With nearly 46,000 stars and over 4,300 contributors, it has become the industry standard for orchestrating complex sequences of tasks, from pulling data out of databases to training AI models.

why it matters: For any company building data-driven products or AI features, Airflow is often the backbone that keeps everything running reliably — making it a critical piece of infrastructure that reduces engineering overhead and accelerates time-to-insight. Its massive adoption signals that data orchestration is now a foundational business need, and teams that implement it early gain a significant operational advantage as their data complexity grows.

46.7k17.7k4.6k contributorsPython

OpenSearch is a free, open-source search and analytics engine that lets you add powerful search functionality to your products — think searching through massive amounts of data, logs, or content instantly. It's the open-source alternative to Elasticsearch, meaning any company can use it without proprietary licensing restrictions.

why it matters: With over 13,000 stars and 2,000+ contributors, OpenSearch has become a serious community-backed alternative to expensive enterprise search tools, giving builders a cost-effective way to add search and data analytics to their products without vendor lock-in. For founders and PMs, this means you can build search-powered features — from site search to security monitoring to business intelligence — on a fully open foundation you control.

13.6k2.9k2.2k contributorsJava

ClickHouse is an open-source database built specifically for analyzing massive amounts of data at lightning speed, returning results in real-time rather than making you wait minutes or hours. Think of it as a supercharged spreadsheet engine that can crunch billions of rows of data almost instantly, making it ideal for dashboards, reports, and any product that needs to show users live insights from large datasets.

why it matters: As user expectations shift toward real-time everything, products that can surface instant insights from data have a significant competitive edge over those with slow, laggy reporting. With nearly 50,000 stars and almost 3,000 contributors, ClickHouse has become a proven, battle-tested foundation that startups and enterprises alike are using to build analytics features without paying the enormous costs of proprietary alternatives like Snowflake or BigQuery.

49.6k8.9k3.1k contributorsC++

form 27-b — subscription

THE TUESDAY BRIEFING

The repos that moved this week, why they matter, and what to watch next. One email. No noise.