Optimus Teslabot Would Be an Edge Computing Beast

Gavin Baker described how in ~3 years (around 2028–2029), a bigger/bulkier iPhone with enough extra DRAM could run a pruned/distilled/quantized version of a frontier model like Gemini 5 (or Grok 4.x / ChatGPT equivalent) at 30–60 tokens per second on-device. This would be free, private, and good enough for most users.

An enhanced future iPhone SoC (Apple Neural Engine or equivalent NPU + CPU/GPU) running a heavily optimized (pruned + quantized to ~4-bit or lower, distilled) version of a Gemini 5-class model.Expected speed: 30–60 tokens/second.

Today, high-end phones already do 10–50+ tokens per second on small-to-medium quantized models (1–7B params, e.g., Gemma-2B, Llama-3 8B variants, Phi models) using llama.cpp, MLX/Core ML, or MediaPipe. Larger pruned frontier models become feasible with better quantization, distillation, and NPU improvements

These are for chat but only limited agentic work.

Tesla will have mass production of the AI5-class chip. The systems would have power in the ~250 W range + much more memory). They taped out April 2026, targeted for vehicles/Optimus ~2027–2028+). Elon Musk has said it will have Blackwell level inference performance for its workloads. Major gains over HW4 will be ~5–8× compute, ~9× memory, ~5× bandwidth, and much higher efficiency.

More memory will let it run with higher precision, larger context windows, or less aggressive pruning without swapping.

Run Fable Class AI for Good Chat and Great Agentic Workload with Gemini 5

On the AI5-class chip (high power + lots of memory)

Memory capacity (100+ GB) + high bandwidth (TB/s) + MoE sparsity (only active experts loaded per token) makes it feasible to run quantized versions (INT4/FP8 or better).
Expected decode speeds running optimized Fable or Grok 5. Roughly 20–100+ tokens/second (or higher with optimizations like speculative decoding, continuous batching, or vLLM-style engines). This is interactive/chat-usable and comparable to or better than dedicated cloud inference setups for frontier models today.