Colibri is a pure-C, zero-dependency inference engine that runs frontier Mixture-of-Experts models like GLM-5.2 (744B parameters) on consumer hardware with only about 25 GB of RAM by streaming routed experts from NVMe disk on demand. It exploits the fact that only ~40B parameters activate per token and only ~11 GB of experts change between tokens, treating VRAM, RAM, and SSD as one staged hierarchy with a per-layer LRU cache and optional pinned hot-store. The engine implements Multi-head Latent Attention with a 57x-compressed KV-cache, AVX2-optimized int8/int4 dot-product kernels at 119 GFLOP/s, and multi-token speculative decoding. Released under Apache 2.0, Colibri lets researchers and builders hold a frontier-class model on their own machine, probe and measure every expert firing in real time, and contribute optimizations. It matters now because it collapses the gap between datacenter-scale MoE inference and a mid-range PC.
codemap is an MIT-licensed project brain for AI coding tools that gives LLMs instant architectural context from your codebase without burning tokens. It generates a fast tree/context view, dependency flow, dependency blast-radius analysis, and a layered handoff format for cross-agent continuation, then exposes everything through a JSON context bundle and an MCP server compatible with Claude Code and Codex. A built-in Codex plugin and community skill registry make it easy to install and share. Developers use codemap to onboard agents to large repos in seconds, keep session continuity across handoffs, and scope the impact of a change before running it.
Type.com is a multiplayer AI workspace where teams collaborate with Claude, Codex and other models in a shared company context. It brings conversations, files, skills, integrations and automations into collaborative Spaces instead of leaving useful work trapped in individual chat windows. Marketing, sales, support and operations teams can tag Type from Slack or email, share access through role-based permissions, and build custom dashboards or internal apps grounded in company knowledge. Type also supports OAuth, MCP and API connections, with granular controls for users and spaces. It is notable now because its Product Hunt launch presents a practical answer to the coordination problem emerging as teams adopt multiple coding and general-purpose agents. The official site confirms a shipped cloud workspace with desktop and mobile access, not merely an agent concept.
Ollama is a local AI platform for running, managing, and sharing open models on your own machine or private infrastructure. It makes it easy to pull models, serve them through an API, and integrate local inference into developer workflows without relying on a fully managed cloud stack. Teams use Ollama for privacy-sensitive assistants, internal tools, offline experimentation, and rapid testing of open-weight models across laptops, workstations, and servers. It is especially useful for developers, operators, and AI builders who want quick setup with less operational overhead. What makes Ollama distinctive is how approachable it is: it packages model runtime, distribution, and deployment into a streamlined experience that helps people get productive with local AI in minutes instead of spending days on configuration.
Ahmad Al-Dahle, the former Meta AI lead now CTO of Airbnb, shares how the company turned 60% of its code over to AI and shipped 80% more features. The real story is how Airbnb collapsed handoffs, built a queryable context graph, and restructured teams around outcomes.
Four days after OpenAI's DevDay, ecosystem signals reveal the real shift: decision models becoming standard infrastructure, platforms tightening permissions, and governance hardifying across the industry.
OpenAI's DevDay 2026 introduced persistent agents (Dots) and GPT-6.1 Sol, signaling a shift from token-based to compute-based pricing. What it means for startups and enterprises.