Voice mode for Claude Code. Claude Code now answers you like a real person
I'll pull the article and HN thread so the paragraphs stay factual and specific.**SKI** (heyski.io) is a desktop app that turns coding agents into a spoken…
By Dillip Chowdary • Aug 06, 2026 • Source: HN Claude/Codex/Fable
I'll pull the article and HN thread so the paragraphs stay factual and specific.**SKI** (heyski.io) is a desktop app that turns coding agents into a spoken loop: you talk, the agent works, and it answers out loud. The HN post frames it as voice mode for **Claude Code** — “Claude Code now answers you like a real person” — and links the product site. On Hacker News the submission sits at **3 points** with **0 comments**. It ships free for Mac and Windows (register once, no card); setup is a single DMG or EXE and claims under two minutes from installer to conversation.
Mechanically it is not dictation that dumps text into a box. Speech recognition and a neural voice both run **on-device**; audio is not uploaded. The floating UI hands text to the agent and speaks replies back, with **full-duplex barge-in** and echo cancellation so you can interrupt mid-sentence on open speakers. The widget and agent exchange work through **plain files inside your project** (inspectable and deletable). Optional **approve-before-send** parks each transcript in an editable bubble until you confirm. On Apple Silicon Macs it docks into the notch (or draws one); on Windows it floats as a pill. A green status indicator shows the agent is listening in real time, not via polling.
Advertisement
Tech Pulse Daily
Get tomorrow's pulse first
Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.
For builders this changes how you drive agents when your hands are on the keyboard, a keyboard, or a second screen. You can bind several repos, each with its own voice, and route speech to the project you mean. A hotkey can attach a screenshot to the next spoken request so the agent sees what you see. Silent mode keeps replies text-only in the pill when you cannot play audio. Mute releases the mic; on Mac, the globe/fn key can toggle mute or act as push-to-talk.
The competitive angle is agent-agnostic surface area, not a Claude-only feature. One skill-style connection covers **Claude Code**, **Cursor**, **Codex**, **Gemini CLI**, **OpenClaw**, and **Windsurf** (with Cursor at project scope and Windsurf marked manual). That puts SKI next to pure STT tools that stop at text-in, and next to any first-party voice work from the agent vendors themselves. Meeting features sit beside the coding loop: local mic-plus-system recording with on-device transcription, and optional agent join into Meet/Teams/Zoom via **AgentCall** (the only paid path; voice and local transcription stay free).
Requirements stay narrow and concrete: **Apple Silicon** (M1 or later) on **macOS 14.4+**, or **Windows 10/11** x64; English only for now, more languages planned; Linux listed as future. The free forever pitch covers the voice loop and local meeting transcriber; you pay only if an agent joins video calls through AgentCall. Watch whether install friction stays near the “one skill per agent” claim in real Claude Code and Cursor projects, whether offline speech quality holds for long coding sessions, and whether agent vendors ship native voice that makes a third-party loop redundant — or whether local audio plus multi-agent routing remains the differentiator.
Advertisement