Latest Announcements

August 1, 2026 🚀 Model & Ecosystem Release

Introducing gemma4-turbo:haven – Uncensored DPO Companion Model & Haven Ecosystem

We are excited to officially announce gemma4-turbo:haven, the flagship DPO-tuned companion variant of G4 Turbo built for the Haven AI Companion ecosystem!

Engineered on top of Gemma 4 E4B with TurboQuant non-linear quantization (IQ4_XS), gemma4-turbo:haven introduces Direct Preference Optimization (DPO) alignment specifically designed for persistent companion roleplay, dynamic state-header tracking, emotional affect modeling, and zero pre-programmed refusal behavior.

Equipped with native vision projector capabilities and 16k context window, gemma4-turbo:haven runs blazingly fast on consumer CPUs without requiring discrete GPU VRAM.

You can pull and run the model immediately on Ollama:

ollama run ssfdre38/gemma4-turbo:haven

Learn more at haven.barrersoftware.com, check out the Ollama Hub Model Page, and explore the Haven AI Companion GitHub organization!

June 29, 2026 🏆 Project Milestone

G4 Turbo surpasses 32k total downloads!

We are thrilled to announce that the G4 Turbo project has officially surpassed 32,700 downloads across the Ollama Hub and Hugging Face!

Our flagship gemma4-turbo variant has climbed to over 25,700 pulls, while the mobile-first gemma4-nano has reached over 6,900 pulls.

As G4 Turbo becomes a major community hub for Google's Gemma 4 models, we want to extend a massive thank you to everyone running, testing, and integrating these models into their workflows!

June 29, 2026 🌐 Network & Ecosystem Release

Announcing Ash Server & Pocket Ash: P2P Compute Grid

We are proud to announce the release of Ash Server and Pocket Ash, expanding the G4 Turbo ecosystem into a fully distributed, private AI network.

Ash Server (.NET 10): A high-performance C# backend that provides local RAG (SQLite vector search), tool calling, and a P2P Compute Grid. By establishing bi-directional WebSocket tunnels, you can pair multiple local servers (like a home NAS, a VPS, and your desktop) to distribute inference load with zero port-forwarding. Check out the code on GitHub ↗.

Pocket Ash (Mobile Client): A standalone mobile application that runs G4 Turbo models offline using LiteRT, or connects securely to your private Ash Server grid via Tailscale, protected by biometric locks. Check out the code on GitHub ↗.

June 17, 2026 🏆 Project Milestone

G4 Turbo surpasses 31k total downloads!

We are thrilled to announce that the G4 Turbo project has officially surpassed 31,000 downloads across the Ollama Hub and Hugging Face!

Our flagship gemma4-turbo variant has climbed to over 24,600 pulls, while the mobile-first gemma4-nano has reached over 6,700 pulls.

This rapid growth is a huge testament to the community's desire for fast, high-performance, and private local AI. Thank you for being a part of this journey!

June 7, 2026 🏆 Project Milestone

G4 Turbo reaches 26.8k total downloads!

We are proud to share that the G4 Turbo project has reached a new milestone: 26,880 downloads across the Ollama Hub and Hugging Face!

Our optimized Gemma 4 builds continue to gain traction as the community embraces high-performance, private, and local AI model options. Thank you for your support!

June 5, 2026 📱 Mobile Project Alpha

Announcing Pocket Ash: Your Local AI Familiar

We are officially beginning the development of Pocket Ash, a standalone mobile application for Android and iOS that brings the full power of G4 Turbo directly to your phone.

Sovereign Privacy: Pocket Ash runs 100% offline using LiteRT. Your conversations and memories never leave your hardware.

Biometric Security: To protect your AI familiar, we have integrated Fingerprint and Face ID locking, ensuring your digital companion is accessible only to you.

June 4, 2026 🚀 New Model Release

G4 Turbo 12B & Nano 12B are here!

We have successfully completed the quantization of the Gemma 4 12B base weights. This release bridges the gap between our lightweight edge models and the massive 31B flagship.

Turbo 12B (IQ4_XS): Optimized for users with 16GB of RAM. This build includes the full Vision Encoder, allowing for high-speed local image analysis and reasoning.

Nano 12B (Q3_K_S): A hyper-compressed 12B model that fits into under 6GB of VRAM, making 12B-tier intelligence accessible to older GPUs and high-end mobile devices for the first time.

June 5, 2026 🏆 Project Milestone

G4 Turbo hits 18.8k pulls on Ollama Hub!

We are thrilled to announce that the gemma4-turbo family has officially surpassed 18,800 pulls on the Ollama Hub. This cements our position as the #1 community-optimized build for Google's Gemma 4.

Our IQ4_XS quantization continues to be the preferred choice for users seeking the perfect balance of performance and quality on consumer CPUs.

Additionally, the gemma4-nano model has seen a massive surge in popularity, hitting 1,117 pulls. This confirms the growing demand for AI that can run on extreme edge hardware like mobile phones and low-RAM devices.

April 15, 2026 🚀 Project Launch

Introducing G4 Turbo: Gemma 4 for Everyone

Today we are launching the G4 Turbo project. Our goal is simple: to make Google's Gemma 4 models run faster and more efficiently on the hardware you already own.

Starting with the e2b and e4b variants, we are releasing optimized weights that offer up to 51% faster inference speeds compared to stock builds.

Follow us on GitHub for real-time code updates.