Researchers from UC Berkeley and UT Austin introduced FreeToken, an edge-native Mixture-of-Experts serving engine that runs models up to GLM-5.2 753B on a single workstation GPU while maintaining interactive inference speeds. FreeToken treats a personal machine as a unified inference platform, dynamically splitting expert computation between GPU and CPU, using semantic-aware caching and elastic memory management to overcome bandwidth limits and traditional MoE engine bottlenecks. Released under Apache-2.0 on GitHub, PyPI, and as desktop apps, FreeToken targets solo developers, startups, and regulated workloads, potentially lowering inference costs and enabling local coding agents, private analysis, and synthetic data generation.
This update represents a notable development in the Ai sector. Organizations and founders tracking this space should evaluate potential strategic and technical implications on their operations.