Catalog of every wiki page. Read this first when answering a query, then drill into the relevant pages. Updated on every ingest.

Sources

Entities

  • TML-Interaction-Small — 276B MoE (12B active); first model strong on both intelligence and interactivity.
  • FD-bench — interactivity benchmark suite (latency, quality, tool use).
  • Moshi — Kyutai’s real-time full-duplex speech-text dialogue model (160ms latency).
  • Mimi (neural audio codec) — streaming neural audio codec (12.5Hz, 1.1kbps) powering Moshi.
  • Helium — 7B text LLM backbone of Moshi, trained from scratch on 2.1T tokens.
  • Voxtral — Mistral’s open-weights audio-understanding models (Mini 4.7B / Small 24.3B); Whisper encoder → adapter → Mistral LLM, audio→text.
  • Qwen3-Omni — Qwen Team’s omni-modal model (30B-A3B MoE); text/image/audio/video in, text/speech out; ~234ms streaming, claims no modality degradation.
  • AuT (Audio Transformer) — Qwen3-Omni’s audio encoder; from-scratch on 20M hrs, 12.5Hz tokens, block-wise window attention for real-time prefill (replaces Whisper).

Concepts

Topics

4 items under this folder.