Star 历史趋势
数据来源: GitHub API · 生成自 Stargazers.cn
README.md

Gemma Translator

This repo was built with the assistance of Google Antigravity and includes code to run an on-device, fully offline voice translator powered by Gemma 4 and LiteRT-LM. This project features a web frontend optimized for small handheld displays (e.g., 480x320) and a Python API server (http.server) that communicates with Gemma. Text-to-speech is powered by Moonshine.

https://github.com/user-attachments/assets/343072ce-dc78-44a7-a783-99312845cabe

Features

  • On-Device Inference: Uses LiteRT-LM to run the gemma4-e2b model entirely locally. No internet required after setup.
  • Voice Interface: Captures microphone audio, processes it, and sends it to the local model.
  • Optimized UI: Retro-terminal styling custom-built for small hardware screens (like Raspberry Pi displays).
  • Unified Startup: One script to launch the LLM server, the Python API, and the React frontend.

Prerequisites

  • Python 3.10+
  • Node.js 18+ (20 LTS recommended) & npm — installed automatically by deploy-pi.sh on Raspberry Pi OS / Debian
  • Linux or macOS

Required Hardware

  • Compute: Raspberry Pi 5 with 8GB RAM
  • Audio Input: Microphone or USB audio capture interface
  • Audio Output: Speaker or headphone output device
  • Display: Display monitor or touchscreen (e.g., 480x320 kiosk display)

Setup Instructions

  1. Make Scripts Executable Ensure the setup, download, start, and deployment scripts have execute permissions:

    chmod +x setup.sh download_model.sh start.sh deploy-pi.sh
  2. Install Dependencies Run the setup script to create a Python virtual environment (venv) and install all required packages:

    ./setup.sh
  3. Download the Model Run the model downloader script to fetch the gemma4-e2b model from Hugging Face and import it into LiteRT-LM:

    ./download_model.sh

Running the Application

Start all services (LiteRT-LM, the Python API server, and the Vite Web UI) in development mode:

./start.sh

To run in production mode (skipping Vite dev server and serving compiled UI assets from frontend/dist/ via backend/server.py on port 3000):

./start.sh --prod

The application will be accessible at:

  • Web UI (Dev): http://localhost:5173
  • Web UI (Prod) / API server: http://localhost:3000
  • LiteRT-LM: http://localhost:9379

Raspberry Pi Appliance Deployment

To deploy as a permanent systemd kiosk service on a Raspberry Pi 5 (8GB):

./deploy-pi.sh

This automated script installs Debian audio/venv packages, sets up the Python environment, builds production UI assets, downloads the LiteRT model, registers the systemd unit from deploy/gemma-translator.service, and configures LXDE GUI autostart (~/.config/lxsession/rpd-x/autostart) to launch Chromium in kiosk mode pointing to http://localhost:3000.

Project Structure

  • frontend/ - React (Vite) web frontend (index.html, src/, styles, and Vite configuration).
  • backend/ - Python API server (server.py and requirements.txt) for Moonshine STT, moonshine-voice TTS, and model proxying.
  • deploy/ - Parameterizable systemd service unit template (gemma-translator.service).
  • stl/ - STL files for 3D printing the hardware case.
  • setup.sh - Automates Python virtual environment creation and dependency installation.
  • download_model.sh - Fetches the required LiteRT model.
  • start.sh - Multi-process launcher supporting --prod and development modes.
  • deploy-pi.sh - One-command Raspberry Pi automated deployment script.

Keyboard Shortcuts

The Gemma Translator supports two keyboard modes. Switch between them anytime from the Settings panel → "Keyboard Mode" dropdown. The choice is remembered across restarts (stored in the browser's localStorage under the key keyboardMode).

The app has two lanes (two people facing each other on the kiosk):

  • Lane 1 / Person 1 — the left/top lane.
  • Lane 2 / Person 2 — the right/bottom lane.

Each lane has a rotating language "revolver" and records speech, which is transcribed (Moonshine STT), translated (Gemma), and spoken back in the other lane's language (moonshine-voice TTS).

Landscape Mode (default) — "active person"

One lane is the active person at a time. The active lane is framed with corner brackets on all four corners. You drive everything from a single set of keys and switch focus with Space.

KeyActionDescription
SpacebarSwitch active personToggles the active lane (Person 1 ⇄ Person 2). Disabled while recording.
ZRecord (push-to-talk)Hold to record the active person; release to transcribe & translate.
← Left ArrowPrevious languageRotates the active person's language backward.
→ Right ArrowNext languageRotates the active person's language forward.

Notes:

  • The active lane shows four-corner brackets; while it is recording, the brackets invert to black along with the lane's color reversal.
  • Best for one-handed / single-operator use.

Vertical Mode — "two-hand" (original mapping)

Each lane has its own dedicated keys — there is no active-person concept and no bracket highlight. Both people can be controlled independently.

KeyActionDescription
ZRecord — Person 1 (push-to-talk)Hold to record Lane 1; release to transcribe & translate.
XRecord — Person 2 (push-to-talk)Hold to record Lane 2; release to transcribe & translate.
← Left ArrowPrevious language — Person 1Rotates Lane 1's language backward.
→ Right ArrowNext language — Person 1Rotates Lane 1's language forward.
− Minus (_)Previous language — Person 2Rotates Lane 2's language backward.
+ Plus (=)Next language — Person 2Rotates Lane 2's language forward.

Notes:

  • No corner-bracket selection highlight in this mode.
  • Best for two operators, each handling their own side.

Common behavior (both modes)

  • Input focus guard: all shortcuts are ignored while focus is on a configuration field (<input>, <textarea>, or <select>) — e.g. when editing the API endpoint or settings.
  • Recording lock: language rotation is blocked while a recording is in progress.
  • Keyboard-driven: recording and language rotation are keyboard-only in the current build; on-screen touch controls are not enabled.

Switching modes

Open Settings (⚙)Keyboard Mode → choose Landscape or Vertical. The change takes effect immediately and persists on the device.

Setting valueMode
landscapeActive-person scheme (Space / Z / ← →) — default
verticalTwo-hand scheme (Z / X / ← → / − +)

Credits

Made by a small team at Google Creative Lab:

Disclaimer

This is not an officially supported Google product. This project is not eligible for the Google Open Source Software Vulnerability Rewards Program.

关于 About

No description, website, or topics provided.

语言 Languages

JavaScript50.9%
Python20.7%
CSS13.9%
Shell13.5%
HTML1.0%

提交活跃度 Commit Activity

代码提交热力图
过去 52 周的开发活跃度
9
Total Commits
峰值: 9次/周
Less
More

核心贡献者 Contributors