Active Prototype Showcase

Real-time voice translation face-to-face

A high-performance Android application built for instant, hands-free spoken conversations across 19 languages. Features automated turn detection, 180° screen flipping, and streaming neural translation.

🎙️
Silero VAD
Automatic turn detection
Kubun Real-Time Voice Translation View
🔄
180° Flip UI
Hands-free conversation flow

Built by Kubun Labs

A family of high-performance, model-augmented Android applications engineered for speed, utility, and modern mobile architecture.

Showcased Project
🌐

Kubun Translate

Real-time face-to-face voice-to-voice translator featuring client-side Silero VAD turn detection, 180° dynamic interface, and NDJSON streaming pipeline.

Kubun Labs
📚

Kubun Notes

Minimalist book note capture app combining on-device ML Kit OCR passage scanning with local Whisper.cpp speech recognition.

Internal Tool
⚡

Kubun Dictate

Self-hosted client-server Whisper flow with GPU backend acceleration for system-wide voice dictation across mobile and desktop.

Engineered for Natural Human Dialogue

Designed from the ground up to remove phone friction in cross-lingual face-to-face interactions.

🎙️

Silero VAD Turn Detection

On-device neural Voice Activity Detection automatically senses speech pauses to trigger end-of-turn translation without touching buttons.

🔄

180° Dual-Facing UI

Dynamic Compose layout rotates speech card 180° for the person sitting across from you, creating an intuitive two-way table companion.

⚡

Sub-Second Speech-to-Text

Audio frames stream to high-throughput Whisper cloud engines, returning clean transcripts with context-aware punctuation.

🌍

19 Supported Languages

Neural Machine Translation with fallback routing ensures accurate idiomatic translation for global travelers and conversations.

🔊

Natural Voice Synthesis

High-fidelity neural text-to-speech outputs natural human inflections so the recipient hears clear, audible voice responses.

📡

Streamed Chunk Pipeline

Edge Workers stream transcript chunks, translation blocks, and base64 audio frames back to Android concurrently to minimize end-to-end latency.

Inside Kubun Translate

A visual tour of the Material 3 Compose interface and audio pipeline.

Translation View

Face-to-Face Conversation Mode

Dual speaker interface with 180° inverted top screen for person B and upright bottom card for person A. Zero tap requirement during active dialogue.

"Ready for voice input • Silero VAD Active • Auto-Detect Language"

Technical Architecture & Stack

Built using Android Clean Architecture with explicit separation across UI, domain, data, and edge infrastructure layers.

Jetpack Compose
Declarative UI with Material 3 components & orientation rotation
Silero VAD
On-device neural silence & end-of-turn speech segmentation
Kotlin Coroutines & Flow
Asynchronous execution & reactive state pipeline
Room DB v2.7
Local conversation persistence excluded from external cloud backup
Cloudflare Workers
TypeScript edge API orchestrator & secret manager proxy
Deepgram STT Engine
High-speed voice-to-text audio transcription
DeepL Neural Translation
Contextual neural translation API across primary target languages
Inworld & OpenAI Voice
Multi-provider Text-to-Speech audio chunk generation

Showcased Engineering Prototype

Kubun Translate is a working Android application prototype built by Andres to demonstrate modern voice architecture, on-device VAD, and edge AI streaming.