A practical field guide to AI models, tools, datasets, and projects. Start with your goal, search the full list, or browse every category.
Every entry keeps the original source close. LifeHubber is a starting point, not an endorsement or safety check, so review current setup, access, terms, and privacy before relying on a project.
Get occasional updates when new AI resources are added
Occasional notes when new AI resources are added. The form below is handled by the mailing-list service, so its own terms apply when you subscribe.
Advertisements
Advertisements
Reader signals
Readers marked these useful
A few entries readers marked useful. Treat this as a signal to explore further, not a ranking or endorsement.
AI ModelsHugging Face
Gemma 4
google/gemma-4
A Google DeepMind Gemma 4 model family collection with public checkpoints including Gemma 4 12B, a dense multimodal model Google describes around local agentic workflows, native audio input, and encoder-free vision/audio handling.
Multimodal models, local agents4 readers found this useful
A newer DeepSeek OCR model release for image/PDF OCR, document-to-Markdown workflows, dynamic resolution, vLLM/Transformers inference, and visual causal flow research.
OCR, document understanding3 readers found this useful
A CLI and channel-routing layer for command-capable agents, with documented paths for web pages, YouTube, RSS, GitHub, Twitter/X, Reddit, Bilibili, Xiaohongshu, Facebook, Instagram, LinkedIn, V2EX, Xueqiu, podcasts, and Exa search, plus doctor checks and safe/dry-run install review.
Agent tooling, web access2 readers found this useful
Recently added projects to open, compare, and check at the source.
AI AgentsHugging Face
Fara1.5-27B
microsoft/Fara1.5-27B
A Microsoft Research 27B multimodal computer-use model for browser tasks that works from screenshots, emits coordinate-grounded actions, and ships with public weights, a Fara harness, and a recommended sandboxed MagenticLite path.
A local dashboard and CLI for running parallel research agents in isolated Git worktrees, tracking experiment branches and evidence, and launching runs across local or remote compute with Claude Code, Codex, or OpenCode.
A local desktop workbench that connects research projects, AI agents, files, code execution, scientific data connectors, artifacts, previews, and an inspectable activity history.
A publicly available 8B vision-language interaction model and deployable system that watches a video stream, decides when to speak or stay quiet, and can delegate harder tasks while keeping the stream in view, with time-aligned data, training materials, and vLLM-based setup paths.
A local-first operator agent runtime for local or configured cloud models, with browser, file, shell, document, memory, task, skill, MCP, HTTP, and Tauri sidecar paths built around a local control loop and on-disk state.
A PrismML family of low-bit multimodal models derived from Qwen3.6-27B, with 1-bit and ternary weights, local inference paths, and a runnable demo repository.
A Cohere Labs 30B-total, 3B-active mixture-of-experts coding model for code generation, agentic software engineering, and terminal tasks, with public BF16 and FP8 Hugging Face weights, Apache 2.0 licensing, and OpenCode or Cohere API try paths listed by official materials.
A Cohere Labs Command A+ model variant with W4A4 quantization, positioned around agentic tool use, multimodal inputs, multilingual work, long context, and Cohere-reported lower hardware requirements.
A newer DeepSeek OCR model release for image/PDF OCR, document-to-Markdown workflows, dynamic resolution, vLLM/Transformers inference, and visual causal flow research.
A Google DeepMind Gemma 4 model family collection with public checkpoints including Gemma 4 12B, a dense multimodal model Google describes around local agentic workflows, native audio input, and encoder-free vision/audio handling.
A Z.ai flagship text-generation model positioned around 1M-token context, long-horizon coding, agentic engineering, public model weights, API access, and local serving paths.
A low-bit on-device translation model from AngelSlim, positioned around 33-language offline translation, GGUF access, Android demo use, and 1.25-bit compression.
A Tencent-Hunyuan multilingual translation model family with 1.8B, 7B, and 30B-A3B variants, 33-language support, GGUF and FP8 options, IFMTBench, training notes, deployment guidance, and Tencent-reported translation results.
A Tencent Hy Team 295B-parameter MoE language model for long-context reasoning, coding, tool use, and agent-style workflows, with public Hy3 and Hy3-FP8 weights, Apache-2.0 licensing, and vLLM or SGLang serving notes.
A Thinking Machines Lab open-weights multimodal model with text, image, and audio input, adjustable reasoning effort, a 1M-token maximum context window, public BF16 and NVFP4 weights, Tinker fine-tuning access, and deployment recipes for major inference frameworks.
An InternLM model series built to reason across scientific paper pages, molecular and material structures, protein sequences, and physical time-series signals, with current Intern-S2 previews, earlier S1 checkpoints, public weights, and deployment guides.
An NVIDIA Isaac GR00T N1.7 vision-language-action model family for humanoid and generalist robot skills, with a 3B model, post-trained variants, GitHub code, inference and fine-tuning notes, LeRobot-format workflow support, and official robotics developer materials.
A publicly available 8B vision-language interaction model and deployable system that watches a video stream, decides when to speak or stay quiet, and can delegate harder tasks while keeping the stream in view, with time-aligned data, training materials, and vLLM-based setup paths.
A Moonshot AI coding-focused agentic model built on Kimi-K2.6, with a 256K context length, image and video input notes, preserve-thinking behavior, multi-step tool-call materials, native INT4 quantization, and deployment paths through vLLM, SGLang, KTransformers, and Moonshot API access.
Compare hosted testing with a 75.2 GB Q4 local route, then inspect Poolside’s published trajectories and documented tool-call limitations before deciding whether this long-context coding model fits.
A ByteDance Research unified multimodal model for image and video understanding, generation, and editing, with model files, demos, inference scripts, Gradio setup, benchmark scripts, and a stated 40GB VRAM inference requirement.
Liquid AI's 230M-parameter LFM2.5 text model for lightweight on-device agentic pipelines, data extraction, and edge inference, with Hugging Face weights, GGUF, ONNX, and MLX variants, 32K context notes, tool-use guidance, and official docs/blog materials.
A Liquid AI hybrid model with 8.3B total parameters, 1.5B active parameters, long-context support, tool-use notes, local formats, and deployment paths for edge and agentic workflows.
An inclusionAI trillion-parameter model for demanding reasoning, coding, tool use, and multi-step agent work, with public weights, publisher-reported evaluations, multi-GPU serving notes, and hosted access paths listed at review time.
A feed-forward 3D foundation model for streaming scene reconstruction, positioned around geometric consistency, long sequences, and efficient real-time inference.
A Meituan LongCat open-weights MoE language model release for coding, long-context, and agentic workflows, with base, FP8, and INT8 Hugging Face model pages, GitHub model-card materials, ModelScope access, and official chat links.
A JetBrains 12B MoE model family for software-engineering AI workflows, with 2.5B active parameters per token, multiple public checkpoints, 131,072-token context notes, tool-use and agent-workflow framing, deployment examples, and a technical report.
A Xiaomi MiMo model family positioned around multimodal understanding, agentic workflows, long-context use, and Pro variants for harder software and tool-heavy tasks.
An OpenBMB 1B-class language model for local assistants, coding agents, tool-use, reasoning, and long-context workflows, with 131K context, Think / No Think modes, standard LlamaForCausalLM architecture, and deployment paths across Transformers, vLLM, SGLang, Docker, GGUF, Ollama, LM Studio, and MLX.
A MiniMax native multimodal model with public ModelScope and Hugging Face pages, 1M context, about 428B total parameters and 23B activated parameters, MiniMax Sparse Attention materials, thinking and non-thinking modes, and deployment paths for Transformers, vLLM, and SGLang.
An OpenMOSS 11B vision-language model family with Realtime, Instruct, and Base checkpoints for continuous video streams, offline multimodal work, and further training, plus public inference and fine-tuning paths.
An NVIDIA 14B text-generation model from the Nemotron-Labs-Diffusion family, focused on switching between autoregressive, diffusion-style parallel decoding, and self-speculation for project-reported decoding efficiency gains.
A NuMind vision-language model for template-guided JSON extraction and document-to-Markdown conversion, with text and image inputs, optional reasoning, multi-page PDF examples, a hosted demo path, full weights, and GGUF variants.
An NVIDIA 4B content-safety model for classifying prompts, optional images, and model responses against standard or custom policies, with a Hugging Face model card, launch post, released dataset, Transformers, vLLM, SGLang, and NVIDIA NIM paths to inspect.
An Ai2 remote-sensing foundation model family for satellite imagery and planetary-scale mapping, with v1.1 Base and BandExtractor models, model weights, training code, a technical report, and Ai2-reported lower compute cost.
A DeepReinforce model family for agentic coding, with public Hugging Face checkpoints from 9B to 397B scale, MIT license tags, GGUF and FP8 variants, and a company post describing self-scaffolding reinforcement learning and project-reported coding-agent benchmark results.
A Qwen language world model for simulating agentic environments, with Apache-2.0 Hugging Face weights, a GitHub repo, AgentWorldBench data, prompts, evaluation/deployment scripts, and domains including MCP, search, terminal, software engineering, Android, web, and OS tasks.
A StepFun multimodal MoE model collection with BF16, FP8, NVFP4, and GGUF variants, 256K context notes, tool-use and agent-workflow framing, and deployment paths across vLLM, SGLang, Transformers, and llama.cpp.
A Google Research time-series forecasting foundation model with a public GitHub repo, TimesFM 2.5 model notes, PyPI package, Hugging Face checkpoints, Apache 2.0 licensing, and related Google BigQuery ML documentation for the supported product path.
A family of vision-language encoders from Google DeepMind, positioned around image-text pretraining, spatial awareness, and general-purpose multimodal applications.
A Microsoft 3D generation model for high-fidelity image-to-3D asset creation, using O-Voxel structured latents, PBR materials, inference code, and training tools.
A Baidu OCR model and code release for one-shot long-horizon document parsing, with public GitHub, Hugging Face, ModelScope, and arXiv materials, Transformers and SGLang examples, and batch image/PDF inference paths.
A small Zyphra mixture-of-experts reasoning model with public weights, 760M active parameters, 8.4B total parameters, deployment notes, and project-reported math and coding evaluations.
A compact 82M-parameter text-to-speech model from hexgrad, with model facts, usage examples, voice materials, samples, a demo Space, and a linked GitHub inference library.
A Liquid AI Hugging Face collection for Japanese-tuned LFMs, grouping LFM2.5-1.2B-JP-202606 and LFM2.5-Audio-1.5B-JP materials for Japanese text, tool use, structured outputs, ASR, TTS, and speech-to-speech workflows.
A robust automatic speech recognition project for messy real-world audio, with code, inference and training paths, Hugging Face weights, a technical report, Voices-in-the-Wild-2M, and project-reported results across difficult acoustic scenarios.
A Xiaomi MiMo speech-recognition model focused on Mandarin, English, Chinese dialects, code-switched speech, noisy audio, songs, and multi-speaker transcription.
A speech and sound generation model family covering TTS, voice design, spoken dialogue, realtime speech, compact speech generation, and MOSS-SoundEffect-v2.0 text-to-audio materials.
An NVIDIA 600M-parameter multilingual streaming automatic speech recognition model for low-latency voice AI and high-throughput transcription, with a Hugging Face model card, NeMo usage paths, language-locales, performance tables, and OpenMDW license terms to inspect.
A Japanese-centric text-to-speech system from SB Intuitions, with Japanese and English generation, style transfer, and zero-shot voice generation support.
An on-device multilingual text-to-speech system built around ONNX Runtime, with local inference, 31-language support, expression tags, model assets, and examples across browser, mobile, desktop, and edge runtimes.
A framework for generating animatable 3D assets from a single image, with mesh, skeleton, and skinning outputs for downstream animation and simulation workflows.
The FastVideo realtime video generation and editing platform, with backend and web UI setup, local GPU, B200, Docker, Modal, readiness checks, and mock-backend workflows.
A local image-generation interface built around prompt-focused SDXL workflows, with Windows downloads, Colab access, inpainting, outpainting, image prompts, and presets.
A Meituan LongCat audio-driven avatar video model for single- and multi-person generation, with audio-text-to-video, audio-image-text-to-video, video continuation, model weights, GitHub quickstart, and project-reported evaluation materials.
An NVIDIA Labs infrastructure codebase for long video generation, with LongLive 2.0 NVFP4 and parallel training/inference support, multi-shot generation, async decoding, model links, docs, configs, and project-reported FPS and VBench results.
Mage-Flow pairs text-to-image generation with instruction-based image editing in one Microsoft 4B model family, with Base, RL-aligned, and four-step Turbo checkpoints plus local and hosted demo paths.
A portrait image-animation framework for live-streaming-style video generation research, with offline and online inference, pretrained weights, a Web UI, and acceleration notes.
An NVIDIA Labs codebase for efficient high-resolution image and video generation, with Sana, Sana-1.5, Sana-Sprint, Sana-Video, training and inference pipelines, model zoo links, ComfyUI and diffusers paths, and newer world-model work.
An agentic video-generation framework for turning ideas, scripts, or longer narratives into planned video workflows, with script generation, storyboards, shot planning, reference selection, consistency checks, and configurable chat, image, and video model providers.
A music-to-dance video model and framework that uses a reference image, a full music track, and a dance-style prompt, with separate global keyframe planning and local refinement stages for longer outputs.
An agent memory layer that distills agent-run learnings into Markdown skill files, with cloud and self-host paths, Python and TypeScript SDKs, dashboard/API surfaces, sandbox and disk tools, and cross-framework reuse.
A Python framework for AI agents and multi-agent workflows, with conversable agents, orchestration patterns, tools, human-in-the-loop flows, code execution options, structured outputs, and an active v1.0 transition note.
A CLI and channel-routing layer for command-capable agents, with documented paths for web pages, YouTube, RSS, GitHub, Twitter/X, Reddit, Bilibili, Xiaohongshu, Facebook, Instagram, LinkedIn, V2EX, Xueqiu, podcasts, and Exa search, plus doctor checks and safe/dry-run install review.
A draft Agentic Resource Discovery specification for publishing and searching agentic resources across federated registries, with ai-catalog.json manifests, REST search, trust metadata, a public Apache-2.0 spec repo, and a Hugging Face Discover implementation for Skills, Spaces, and MCP servers.
A curated library of medical research agent skills designed to support evidence review, protocol design, data analysis, and academic writing workflows.
A local-first AI workspace for chatting with documents, using agents, connecting local or cloud model providers, adding MCP tools, building agent flows, and running as desktop, self-hosted, or hosted software.
A local-first operator agent runtime for local or configured cloud models, with browser, file, shell, document, memory, task, skill, MCP, HTTP, and Tauri sidecar paths built around a local control loop and on-disk state.
A browser automation framework for AI agents that can navigate websites, click elements, type into pages, use custom tools, and run browser tasks through code or CLI workflows.
A local codebase knowledge graph for coding agents, with MCP-style tools, symbol relationships, call graphs, framework-aware routes, auto-sync, multi-agent setup paths, and project-reported savings in token use and tool calls.
An agent tool-integration layer with Python and TypeScript SDKs, toolkits, authentication, sessions, triggers, tool search, and workbench features for connecting agents to external services.
A frontend stack for building agent-native applications with chat UI, generative UI, shared state, backend tool rendering, human-in-the-loop workflows, and React or Angular app paths.
A computer-use agent stack with Cua Driver for background desktop control, agent-ready sandboxes, CuaBot, Cua-Bench, Lume, SDKs, MCP support, and model integrations.
An agent-native personalized tutoring system with tutoring workflows, persistent memory, a web app, CLI access, and a broader learning-support architecture.
A visual platform for building agentic workflows and AI applications with workflow and chatflow builders, model-provider connections, RAG pipelines, tools, APIs, logs, and cloud or self-hosted paths.
A small skills repository for design engineers, with agent guidance around UI polish, animation decisions, strict motion review, and animation vocabulary for prompting designers or AI coding agents.
A Sentient AGI toolkit for benchmark-driven coding-agent skill discovery and improvement, with support for Claude Code, Codex CLI, OpenCode, OpenHands, Goose, Harbor, local/Docker/remote runs, logs, diffs, and reusable evolved skill folders.
A Microsoft Research 27B multimodal computer-use model for browser tasks that works from screenshots, emits coordinate-grounded actions, and ships with public weights, a Fara harness, and a recommended sandboxed MagenticLite path.
An independent Chrome extension experiment for running an on-device browser agent with Transformers.js, WebGPU, Gemma 4, page RAG, tab tools, and semantic history search.
A Google Colab command-line interface for connecting local terminal and agent workflows to remote Colab runtimes, with CPU/GPU/TPU provisioning, local script and notebook execution, artifact and log retrieval, REPL or console access, and Linux/macOS-only support at launch.
A public Agent Skills repository for Google products and technologies, including Google Cloud, with installable skills for Gemini API in Agent Platform, cloud basics, onboarding, authentication, observability, and well-architected guidance.
A local-first context compression layer for AI agents and LLM apps, with Python and TypeScript library paths, proxy mode, MCP tools, agent wrappers, reversible cached originals, cross-agent memory, failure-learning utilities, and project-reported token-savings benchmarks.
A Hugging Face Space for reading Claude Code session JSONL traces, reconstructing sessions in plain English, surfacing production/config/secret-related activity, showing tool and token usage, and answering trace questions with turn references.
A Nous Research agent runtime with CLI and TUI use, persistent memory, reusable skills, scheduled automations, messaging gateways, subagents, tool configuration, provider switching, and multiple terminal backends.
A beta agent-workspace environment for recurring AI work-streams, with living workspaces, memory, history, files, apps, dashboards, runtime state, and sub-agent coordination.
An H Company computer-use model family for web, desktop, and mobile automation, with 0.8B, 4B, 9B, and 35B-A3B sizes plus FP8, NVFP4, and Q4 GGUF checkpoint paths for local or edge-oriented deployment.
Hugging Face describes hf as the command-line entrypoint for working with the Hub, with agent-mode output, predictable command help, skills support, and workflows for Hub tasks such as search, metadata, downloads, uploads, repos, jobs, collections, and inference.
A Hugging Face GitHub-native AI code reviewer for pull requests, with repository-owned review rules, GitHub Action, GitHub App, and staged web app modes, human editing before publication, and OpenAI-compatible model-provider options.
A Python diagnostic for running fixed behavioral and governance inspections against language models and AI agents, with provider adapters, repeatable suites, cross-provider judging, manifests, scorecards, and CLI, CI, plugin, or agent-skill paths.
A JetBrains Kotlin and Java framework for building AI agents with tools, graph workflows, memory, RAG, MCP, A2A, Agent Client Protocol, tracing, Spring Boot, Ktor, and multiple LLM providers.
A local context layer for data agents that connects SQL warehouse context, semantic definitions, wiki knowledge, CLI search, and MCP tools so agents can reuse approved data context.
A Codex/Desktop-oriented agent skill for building last-30-days topic briefs across Reddit, Hacker News, Google Trends, YouTube, podcasts, GitHub, web search, news, and blogs, with installable skill/plugin files, source-attribution rules, tests, and recent changelog activity.
A realtime framework for voice, video, and physical AI agents, with Python and Node.js paths, LiveKit room participants, WebRTC clients, telephony support, tools, testing, and deployment options.
A cross-platform desktop app for turning documents into an LLM-maintained wiki, with source traceability, graph search, optional vector retrieval, local API access, and a companion agent skill.
A TypeScript framework for building AI agents and applications with model routing, workflows, human-in-the-loop steps, memory, tools, MCP servers, evals, and observability.
An open-source knowledge-base agent platform for building RAG-backed Q&A assistants and agent workflows, with document upload and web-source collection, vector/full-text/hybrid retrieval, model-provider connections, MCP/tool use, web embedding, Docker/offline deployment paths, and GPL-3.0 licensing.
A memory layer for AI agents and assistants, with library, self-hosted, platform, SDK, CLI, cookbook, evaluation, and integration paths for persistent context.
A Rust-based memory layer for AI agents that keeps documents, embeddings, search indexes, metadata, and recovery state in one portable .mv2 file, with CLI, Python, Node.js, and framework integration paths.
A curated library of portable agent skills for designers and builders using Codex, Claude, Cursor, Aura Build, Lovable, and similar AI coding agents, with workflows for design prompts, landing pages, screenshots, WebGL, animation, frontend systems, and reusable skill folders.
A Cosmic Stack personal AI agent runtime with permission modes, local structured memory, editable personality files, token budgets, provider choices, skills, scheduled work, background service operation, and terminal, web, or messaging channels.
A Microsoft framework for building AI agents and multi-agent workflows across Python and .NET, with agents, graph workflows, middleware, MCP integrations, context providers, observability, and provider choices.
A Microsoft toolkit for adding policy enforcement, identity, sandboxing, audit records, reliability tooling, and the Agent Control Specification runtime layer around AI agents, with Python, TypeScript, .NET, Rust, and Go surfaces.
A Microsoft Responsible AI evaluation harness for AI agents and LLM apps that turns natural-language requirements into generated test scenarios, runs them against models or traced agents, and writes local artifacts for review.
An agent and product integration platform for auth, OAuth, API proxying, TypeScript functions, MCP tool calling, observability, and external API access.
A lightweight personal AI agent project with packaged WebUI, goal tracking, Python SDK runtime controls, automation management, provider and channel integrations, and durability work highlighted in its v0.2.2 release.
An NVIDIA AI Blueprint for agentic research workflows, with shallow and deep research modes, citation-backed answers, YAML-configured tools and agents, CLI/web/async job paths, evaluation materials, and deployment assets.
NVIDIA's public catalog of agent skills, framed around NVIDIA-verified skills, skill cards, scanning, signing, product-owned source repositories, and compatibility with the Agent Skills specification.
A computer-use MCP service for AI agents and MCP clients, with macOS, Linux, and Windows paths, Codex, Claude Code, Gemini CLI, opencode setup commands, a Codex skill path, command-call examples, and local validation commands.
A TypeScript-native framework for coordinating multi-agent work from a goal into a task DAG, with auto task decomposition, parallel execution, plan preview and replay, human approval hooks, MCP tools, provider routing, observability, and sandboxed filesystem defaults.
A local desktop workbench that connects research projects, AI agents, files, code execution, scientific data connectors, artifacts, previews, and an inspectable activity history.
A Hugging Face framework for creating, deploying, and using isolated execution environments for agentic RL training, with Gymnasium-style reset, step, and state APIs, client/server environment access, HTTP and WebSocket protocols, Docker packaging, MCP compatibility, and examples for environment builders.
An OpenHands developer control center for coding agents and automations, with official paths for OpenHands, Claude Code, Codex, Gemini, ACP-compatible agents, local/Docker/VM/cloud backends, Agent Canvas UI, Agent Server, prebuilt automations, and a project transition toward Agent Canvas and the Software Agent SDK.
A local dashboard and CLI for running parallel research agents in isolated Git worktrees, tracking experiment branches and evidence, and launching runs across local or remote compute with Claude Code, Codex, or OpenCode.
A Mac-native AI agent harness for Apple Silicon, with local and cloud model options, agents, memory, tools, cryptographic identity, MCP and local API paths, privacy-filter materials, plugin routes, and optional macOS 26 sandbox features.
A JavaScript in-page GUI agent for adding natural-language control to web interfaces, with BYOK model setup, text-based DOM interaction, demo and npm paths, plus optional Chrome extension and MCP server support.
A Python framework and ecosystem for real-time voice and multimodal AI agents, with audio/video pipelines, transports, client SDKs, structured flows, and subagent support.
A practical RAG and agent-context platform for document ingestion, chunking, retrieval, citations, knowledge workflows, and self-hosted AI applications.
A Microsoft research project for optimizing reusable natural-language skills for frozen LLM agents, with trajectory-driven edits, validation-gated updates, benchmark configs, training and evaluation scripts, WebUI monitoring, and deployable best_skill.md artifacts.
An NVIDIA public scanner for AI agent skills, with CLI scans for Git repositories, URLs, zip files, directories, and single files; static analysis, optional LLM semantic review, OSV dependency lookups, and terminal, JSON, Markdown, or SARIF reports.
A local memory plugin for AI agents, with symbolic short-term memory, layered long-term memory, SQLite defaults, OpenClaw integration, Hermes support, and project-reported benchmark results.
A ByteDance GUI-agent desktop app and multimodal agent stack for local or remote computer and browser operation, with vision-language model control, CLI/Web UI paths, and MCP-oriented tooling.
A codebase and knowledge-base graph tool for AI coding environments, with plugin paths for Claude Code, Codex, Cursor, Copilot, Gemini CLI, and others, plus search, chat, tours, diff impact views, and an interactive dashboard.
A self-hostable AI answering engine for private search-style workflows, with local and cloud model providers, SearxNG-backed web search, cited sources, file uploads, and Docker setup paths.
A Hugging Face robotics library for robot-learning workflows, with models, datasets, robot interfaces, training, evaluation, deployment, and v0.6.0 updates across world models, reward models, benchmarks, dataset tooling, and cloud training paths.
A public humanoid robotics workspace around a low-cost 3D-printed robot-learning platform, with hardware files, assembly documentation, runtime tools, simulation/model assets, identification tooling, and training environments.
A browser-based multi-track video editor with local workspace files, on-device transcription and captions, scene analysis, local voice and music tools, effects, subtitles, and in-browser export through modern Chromium APIs.
A local agentic HTML editor that uses existing coding-agent CLI sessions to turn Markdown, data, and notes into exportable HTML, PNGs, decks, social cards, data reports, and web prototypes with skill templates and sandboxed preview.
An open source desktop AI assistant for meeting notes, dictation, projects, and local agent sessions, with a Tauri app, June API backend, local app state by default, macOS and Windows release paths, and TEE verification notes for the backend.
A local-first AI meeting assistant with public source, desktop release assets for Windows and macOS, Linux source-build docs, live transcription, meeting summaries, local Whisper or Parakeet transcription, local or external LLM summary options, optional analytics, and an MIT license.
A Xiaomi MiMo terminal-native AI coding assistant built on OpenCode, with public source, docs, install paths, build/plan/compose agents, persistent memory, task tracking, subagents, provider configuration, and release assets.
A local-first AI design workspace that connects coding-agent CLIs to prototypes, decks, media outputs, design systems, sandboxed previews, and export workflows.
An agent-native slide framework for building React-based decks with coding agents, browser preview, comments, assets, present mode, and HTML/PDF export.
An open-source collaborative design platform with an official MCP server that lets AI agents inspect and modify Penpot files, components, tokens, styles, layouts, and assets through hosted or local connection paths.
A self-hosted AI accounting app for receipts, invoices, PDFs, and transactions, with custom prompts, fields, categories, provider or local LLM options, Docker/PostgreSQL setup, exports, and an early-development warning.
A lightweight AI-native terminal and developer environment built with Tauri, Rust, and React, with a terminal, code editor, file explorer, web preview, AI side panel, BYOK providers, local model support, and approval-style file tools.
An agentic development environment born out of the terminal, with built-in coding-agent workflows and support for bringing external CLI agents into developer work.
An Apple repository for building on-device AI around Core AI, with export recipes for supported open-source models, Python primitives for PyTorch authoring, Swift runtime utilities for macOS and iOS apps, a model catalog, CLI paths, and coding-agent skills.
A curated collection of DESIGN.md example files inspired by public websites, intended to help AI coding agents understand visual systems, design tokens, layout rules, and UI guardrails.
An incremental data engine for keeping AI-agent and LLM-app context fresh, with Python-native pipelines, delta-only processing, lineage, connectors, and targets for vector, graph, relational, and warehouse stores.
TencentCloud sandbox infrastructure for AI agents, with KVM MicroVM execution, E2B-compatible code workflows, templates and snapshots, and a v0.5.0 release covering AutoPause/AutoResume, ARM64 support, a TencentCloud Terraform path, and network-hardening notes.
A synthetic data generation framework for creating structured datasets from scratch or seed data, with dependency-aware generation, validation, and quality scoring.
A format specification and CLI toolkit for describing a design system to coding agents, positioned around persistent visual guidance, linting, and token-level design workflows.
A Zig-based headless browser for AI agents and automation, with JavaScript and DOM execution, CDP connections for Puppeteer, Playwright, and Chromedp, local and cloud run paths, a native agent that can save reusable PandaScripts, and MCP support.
A smart model router for personal AI agents, positioned around cost-aware request routing, fallbacks, provider control, and self-hosted agent workflows.
A focused Python and Streamlit workflow for using Ollama vision models to extract text and structured output from images or PDFs, with preprocessing, batch runs, custom prompts, and multiple output formats.
An Ai2 toolkit for converting PDFs, PNGs, and JPEG document images into clean Markdown or text for downstream AI workflows, with local GPU, remote vLLM/OpenAI-compatible server, Docker, benchmark, model-card, and demo paths.
A privacy-filtering model and local toolkit for detecting and masking personally identifiable information in text, positioned around high-throughput sanitization workflows.
Alibaba sandbox runtime infrastructure for AI applications, with multi-language SDKs, unified sandbox APIs, Docker and Kubernetes runtimes, CLI and MCP paths, code-interpreter support, browser and desktop examples, and network-control features for agent workloads.
A Hugging Face toolkit for exporting, quantizing, compressing, and running Hub models through OpenVINO on Intel CPUs, Arc GPUs, and Core Ultra NPUs, with v2.0 making OpenVINO and NNCF the default path and release notes covering newer text, vision-language, speech, video, and diffusion model support.
A document AI toolkit for OCR, document parsing, structured Markdown and JSON outputs, PaddleOCR-VL document parsing, PP-StructureV3 conversion, PP-OCRv6 scene OCR, and workflows that feed RAG or agent systems.
A Datalab document OCR and analysis toolkit for OCR, layout analysis, reading order, table recognition, math-aware output, vLLM or llama.cpp backends, and structured document results.
An experimental graph-first programming language for agents, where reviewable .0 source stays the durable artifact while compiler-derived ProgramGraph facts, graph hashes, node IDs, diagnostics, checked edits, and version-matched skills give coding agents a more structured interface.
A manually curated benchmark for general reasoning in LLMs, designed around high difficulty, broad task diversity, K-12-scope knowledge, and hybrid scoring.
A benchmark for evaluating latent action representations, with pipelines for action semantics, robotic control regression, and broader vision-to-action alignment.
An OpenAI evaluation-data release for studying monitorability, with public eval splits, prompt templates, dataset mappings, and metric code from the Monitoring Monitorability paper.
A public results dataset for the Open Agent Leaderboard, with agent, model, benchmark, score, completion, error, action-count, and cost fields linked to the Exgentic evaluation framework and paper.
A terminal-agent benchmark for evaluating AI agents on hard containerized command-line tasks, with Harbor run commands, task-level registry pages, GitHub and Hugging Face materials, docs, and paper links.
Keep the thread going with AI Guides for decision habits for messy AI choices, AI Access for free and low-cost ways to compare AI model access, AI Ballot for a clearer view of what readers are leaning toward.