今天,我们宣布 Antigravity SDK 支持跨多种本地模型和执行选项的本地工作流,并初步支持使用 Google AI Edge 的 LiteRT 运行 Gemma 4 26B A4B。
Antigravity SDK 使开发人员能够利用驱动 Google Antigravity 的相同智能体(agentic)能力进行构建。借助这一新支持,您可以通过本地模型完全离线地启用智能体辅助功能。我们针对 LiteRT 和 Gemma 4 26B 优化了此工作流,高效利用本地 GPU 和内存,从而进一步放大您的本地机器所能提供的性能!
为什么要在本地运行智能体?
本地模型执行为智能体体验提供了多项优势:
成本效益:执行本地智能体工作流,无需支付 API 费用或受速率限制。
隐私保护:将您的代码和请求完全保留在本地计算机上,非常适合需要应对严格数据隐私要求或合规受限企业环境的开发人员。
离线韧性:即使在无法保证稳定网络连接的环境中,也能无缝执行智能体工作流。
混合工作流:将令牌高效的本地流程与基于云的流程相结合,以最大化效率,同时在需要时保留对更大、更强大模型的访问权限。
以下是您的入门指南:(我们建议配备 >24GB VRAM 或统一内存的机器)。
创建虚拟环境:
python3 -m venv .venv
source .venv/bin/activate
Python
安装 Antigravity SDK 和 LiteRT-LM,并下载 Gemma 4 26B A4B:
pip install google-antigravity litert-lm
litert-lm import \
--from-huggingface-repo=litert-community/gemma-4-26B-A4B-it-litert-lm \
gemma-4-26B-A4B-it-gpu.litertlm \
gemma4-26b
Python
在您的目录中,创建一个名为 agy_sample.py 的文件。将以下内容粘贴到其中。
import asyncio
import os
from google.antigravity import Agent, LiteRTAgentConfig
from google.antigravity.hooks import policy
MODEL_PATH = os.path.expanduser("~/.litert-lm/models/gemma4-26b/model.litertlm")
async def main():
print(f"Using local LiteRT model: {MODEL_PATH}. Please wait for local inference to complete. This could take several minutes.")
config = LiteRTAgentConfig(model_path=MODEL_PATH).lightweight()
async with Agent(config) as agent:
response = await agent.chat("What files are in the current directory?")
async for token in response:
print(token, end="", flush=True)
if name == "main":
asyncio.run(main())
Python
混合编排:云架构师与设备端工作力量的结合
在许多情况下,我们看到“架构师-构建者”模式是将云模型规模优势与本地模型优势相结合的好方法。在下面的混合演示视频中(使用更新后的 Antigravity SDK 构建),云架构师(Gemini 3.8 Flash)担任规划者和指挥者,而由 Gemma 4 26B 实例组成的本地群体则在设备端完全承担繁重任务。
当被要求审计和修补三个易受攻击的模块(auth.py、billing.py 和 database.py)时,该工作流保持了严格的数据隐私,并让我们能够充分利用令牌利用率:
无代码上传:Gemini 3.8 Flash 仅基于文件名和任务描述来规划策略并分解工作——仅消耗 95 个云令牌,且源代码从未离开过机器。
自主本地试炼:本地 Gemma 4 26B 模型接管本地 GPU,执行对抗性审计循环:复现安全漏洞、编写候选修复方案、审查补丁以及针对回归测试套件进行验证。
巨大的成本与隐私收益:在此录制的运行中,97.2% 的所有令牌(3,322 个令牌)在本地和离线状态下运行,无需调用云 API,同时交付完全验证通过的绿色补丁,并将专有代码完全安全地保留在设备端。
在此查看示例项目,以运行内置的3文件闯关测试,或将其指向您自己的Python模块和测试套件。
抱歉,您的浏览器不支持此视频播放
免费Token本地实用工具:CLI资源监控器
Antigravity SDK结合Gemma 4 26B A4B在构建实用的系统工具方面表现出色。在此示例中,代理构建了一个在终端中运行的实时更新的资源监控器。只需一个提示,代理便自主编写了一个使用psutil和rich库来追踪CPU和内存用量的Python脚本,生成必要的requirements.txt文件,甚至测试生成的代码以确保其正常工作——所有操作均在您的本地机器上运行,并使用Gemma 4 26B模型。
抱歉,您的浏览器不支持此视频播放
使用Gemma 4 26B在设备上生成的基于Python的CLI资源监控工具
import asyncio
import os
from google.antigravity import Agent, LiteRTAgentConfig
from google.antigravity.hooks import policy
PROMPT = "使用psutil和rich库构建一个命令行界面工具,显示实时更新的终端仪表板。它应显示CPU使用情况、内存消耗以及按内存占用排序的前5个最耗内存进程的表格。将脚本保存为'monitor.py'并创建'requirements.txt'文件。测试其是否正常工作。"
MODEL_PATH = os.path.expanduser("~/.litert-lm/models/gemma4-26b/model.litertlm")
WORKING_DIR = os.path.expanduser("~/agy-test")
os.makedirs(WORKING_DIR, exist_ok=True)
os.chdir(WORKING_DIR)
async def main():
print(f"使用本地LiteRT模型: {MODEL_PATH}。请等待本地推理完成。这可能需要几分钟时间。")
config = LiteRTAgentConfig(
model_path=MODEL_PATH,
workspaces=[WORKING_DIR],
policies=[policy.allow_all()],
).lightweight()
async with Agent(config) as agent:
response = await agent.chat(PROMPT)
async for token in response:
print(token, end="", flush=True)
if name == "main":
asyncio.run(main())
Python
Antigravity SDK还通过LocalOpenAIAgentConfig为任何与OpenAI兼容的服务器(如Ollama、LM Studio或vLLM)提供无缝即插即用支持。这使您能够灵活地尝试不同的本地推理后端,同时保持代理编排、工具和工作流程完全不变。
请查看Antigravity Python SDK README中的说明以开始使用本地AI,并了解更多关于如何使用LiteRT在边缘高效运行模型的信息。请在Antigravity Python SDK GitHub Issue Tracker上分享您的反馈和功能请求。我们期待看到您构建的内容!
致谢:Abhi Patel, Ander Dobo, Ben Miles, Cormac Brick, Ian Ballantyne, Jingxiao Zheng, Jonathan Reay, Kimish Patel, Lu Wang, Marissa Ikonomidis, Matthias Grundmann, Olivier Lacombe, Omar Sanseviero, Rishika Sinha, Rody Davis, Taylor Mullen, Tyler Mullen, Wai Hon Law, Xiaoming Hu, Xu Chen, Yu-hui Chen
Today, we’re announcing that the Antigravity SDK supports local workflows across a wide range of local models and execution options, featuring initial support for Gemma 4 26B A4B using Google AI Edge ’s LiteRT .
The Antigravity SDK enables developers to build with the same agentic capabilities that power Google Antigravity . With this new support you can enable agentic assistance via local models completely offline. We’ve optimized this workflow for LiteRT and Gemma 4 26B, efficiently using the local GPU and RAM in order to further amplify what your local machine is capable of delivering!
Why run agents locally?
Local model execution offers several advantages for agentic experiences:
Cost efficiency: Execute local agentic workflows without API costs or rate limits.
Privacy: Keep your code and requests entirely on your local machine, ideal for developers navigating strict data privacy requirements or compliance-restricted corporate environments.
Offline resiliency: Execute your agentic workflows seamlessly, even in environments where a consistent or stable internet connection is unavailable.
Hybrid workflows: Combine token-efficient local processes with cloud-based ones to maximize efficiency while retaining access to larger, more powerful models when needed.
Here is how you can get started: ( We recommended a machine with >24GB VRAM or unified memory).
Create a virtual environment:
python3 -m venv .venv
source .venv/bin/activate
Python
Install Antigravity SDK and LiteRT-LM, and download Gemma 4 26B A4B:
pip install google-antigravity litert-lm
litert-lm import \
--from-huggingface-repo=litert-community/gemma-4-26B-A4B-it-litert-lm \
gemma-4-26B-A4B-it-gpu.litertlm \
gemma4-26b
Python
In your directory, create a file called agy_sample.py . Paste the following contents into it.
import asyncio
import os
from google.antigravity import Agent, LiteRTAgentConfig
from google.antigravity.hooks import policy
MODEL_PATH = os.path.expanduser("~/.litert-lm/models/gemma4-26b/model.litertlm")
async def main():
print(f"Using local LiteRT model: {MODEL_PATH}. Please wait for local inference to complete. This could take several minutes.")
config = LiteRTAgentConfig(model_path=MODEL_PATH).lightweight()
async with Agent(config) as agent:
response = await agent.chat("What files are in the current directory?")
async for token in response:
print(token, end="", flush=True)
if name == "main":
asyncio.run(main())
Python
Hybrid Orchestration: Cloud Architect Meets On-Device Workforce
In many cases we see that an Architect-Builder pattern is a great way of combining cloud model scale with local model advantages. In the hybrid demo video below, built with the updated Antigravity SDK, a cloud architect (Gemini 3.8 Flash) acts as the planner and conductor, while a local swarm of Gemma 4 26B instances handles the heavy lifting entirely on-device.
When tasked with auditing and patching three vulnerable modules (auth.py, billing.py, and database.py), the workflow maintains strict data privacy and allows us to make the most of our token utilization:
No code uploaded: Gemini 3.8 Flash plans the strategy and decomposes the work based purely on filenames and task descriptions - spending just 95 cloud tokens without any source code ever leaving the machine.
Autonomous local gauntlet: Local Gemma 4 26B models take over on the local GPU to execute an adversarial audit loop: reproducing security vulnerabilities, authoring candidate fixes, critiquing patches, and validating against regression test suites.
Massive cost and privacy wins: In this recorded run, 97.2% of all tokens (3,322 tokens) run locally and offline without calling a cloud API, delivering fully verified, green patches while keeping proprietary code completely secure on-device.
Check out the example project here to run the built-in 3-file gauntlet or point it at your own Python modules and test suite.
Sorry, your browser doesn't support playback for this video
Token Free Local Utilities: CLI Resource Monitor
The Antigravity SDK with Gemma 4 26B A4B excels at building practical system utilities. In this example, the agent built a live-updating resource monitor that runs in the terminal. Given a single prompt, the agent autonomously writes a Python script that uses the psutil and rich libraries to track CPU and memory usage, generates the necessary requirements.txt file, and even tests the resulting code to ensure it works - all running entirely on your local machine and using Gemma 4 26B.
Sorry, your browser doesn't support playback for this video
A Python based CLI Resource Monitor tool generated on-device with Gemma 4 26B
import asyncio
import os
from google.antigravity import Agent, LiteRTAgentConfig
from google.antigravity.hooks import policy
PROMPT = "Build a command-line interface tool using the psutil and rich libraries that displays a live-updating terminal dashboard. It should show CPU usage, memory consumption, and a sorted table of the top 5 most memory-intensive processes. Save the script as 'monitor.py' and create a 'requirements.txt' file. Test that it works."
MODEL_PATH = os.path.expanduser("~/.litert-lm/models/gemma4-26b/model.litertlm")
WORKING_DIR = os.path.expanduser("~/agy-test")
os.makedirs(WORKING_DIR, exist_ok=True)
os.chdir(WORKING_DIR)
async def main():
print(f"Using local LiteRT model: {MODEL_PATH}. Please wait for local inference to complete. This could take several minutes.")
config = LiteRTAgentConfig(
model_path=MODEL_PATH,
workspaces=[WORKING_DIR],
policies=[policy.allow_all()],
).lightweight()
async with Agent(config) as agent:
response = await agent.chat(PROMPT)
async for token in response:
print(token, end="", flush=True)
if name == "main":
asyncio.run(main())
Python
The Antigravity SDK also offers seamless, plug-and-play support for any OpenAI-compatible server such as Ollama, LM Studio, or vLLM via LocalOpenAIAgentConfig . This gives you the flexibility to experiment with different local inference backends while keeping your agent orchestration, tools, and workflows completely unchanged.
Get started with local AI by checking the instructions on the Antigravity Python SDK README , and learn more about how you can run models efficiently on the edge using LiteRT . Please share your feedback and feature requests on the Antigravity Python SDK GitHub Issue Tracker. We look forward to seeing what you build!
Acknowledgements: Abhi Patel, Ander Dobo, Ben Miles, Cormac Brick, Ian Ballantyne, Jingxiao Zheng, Jonathan Reay, Kimish Patel, Lu Wang, Marissa Ikonomidis, Matthias Grundmann, Olivier Lacombe, Omar Sanseviero, Rishika Sinha, Rody Davis, Taylor Mullen, Tyler Mullen, Wai Hon Law, Xiaoming Hu, Xu Chen, Yu-hui Chen
| 刊期 | 得分 | 排名 | 结果 |
|---|---|---|---|
| 2026-10-02 | 8.4 | 46 | 入选 |
| 2026-09-24 | 8.4 | 59 | 入选 |