GLOBAL WORK • Worldwide Delivery · Customers & Projects worldwide
🌐 100% Remote-First ⚡ 24 / 7 Deployment

IA agents with voice: Whisper + Kokoro TTS in production

18 August 2026 · Euskolabs

In euskolabs we have Charly, a hybrid voice assistant operating 24 / 7 on Telegram, web and API. This is the technical guide to how we built it using Faster-Whisper for voice recognition and Kokoro TTS for synthesis, with local offline fallback.

System architecture

Audio entrada → Whisper (STT) → LLM (razonamiento) → Kokoro (TTS) → Audio salida
                          ↓
                    Ollama (fallback local)
                          ↓
                    Gemini (cloud)

The flow is simple: the user speaks, Whisper transcribe, the LLM generates the answer, Kokoro synthesizes it to voice. The total back and forth latency is 2-4 seconds in standard hardware.

Faster-Whisper - Voice recognition

Faster-Whisper is an optimized reimplementation of Whisper using CTranslate2. It's 4x faster than the original Whisper with the same accuracy.

from faster_whisper import WhisperModel

model = WhisperModel("base", device="cpu", compute_type="int8")

def transcribe(audio_path):
    segments, info = model.transcribe(audio_path, language="es")
    text = " ".join([s.text for s in segments])
    return text.strip()

We use the model base in CPU because it offers the best speed / precision balance. For GPU environments, large-v3 gives almost perfect precision.

Kokoro TTS - Natural voice synthesis

Kokoro is an open source TTS engine with impressive natural quality. It supports multiple voices and languages.

import kokoro

# Inicializar con voz por defecto
tts = kokoro.KokoroTTS(model_path="kokoro.pt", voice="em_alex")

def speak(text, output_path="/tmp/response.wav"):
    tts.generate(text, output_path)
    return output_path

The voice em_alex is the one Charly uses in production - natural, clear, with good rhythm. The 10-second audio generation takes less than 2 seconds in CPU.

LLM with fallback cloud → local

The agent's brain uses double mode: Gemini in the cloud for maximum capacity and Olama with local LLAMA for privacy and offline operation.

import google.generativeai as genai
import ollama

def generate_response(prompt, context=""):
    try:
        # Cloud first — máxima capacidad
        model = genai.GenerativeModel("gemini-pro")
        response = model.generate_content(f"{context}\n\nUser: {prompt}")
        return response.text
    except Exception:
        # Fallback local — privacidad y offline
        response = ollama.chat(
            model="llama3",
            messages=[{"role": "user", "content": f"{context}\n\n{prompt}"}]
        )
        return response["message"]["content"]

The fallback is automatic: if the Gemini API fails (rate limit, network fall, etc.), the system uses local Ollama without any interruption visible to the user.

Integration with n8n

The agent connects to business processes via n8n webhooks:

import httpx

async def execute_workflow(action, params):
    webhook_url = f"http://n8n:5678/webhook/{action}"
    async with httpx.AsyncClient() as client:
        response = await client.post(webhook_url, json=params)
        return response.json()

This allows the agent to run real actions: create invoices, send emails, update CRM, etc.

Actual deployment

Charly runs in a VPS with Docker. The models of Whisper and Kokoro are loaded in memory when starting. Olama runs as a separate service. The total latency:

PhaseTime
Whisper (transcription)~ 0.5s
LLM (response generation)~ 1-2s
Kokoro (voice synthesis)~ 1-2s
Total back and forth~ 2-4s

Conclusions

You want your own AI Agent with a voice? Request a diagnosis in euskolabs. com.

← Back to the blog

↑