🪴 Prabhav's Log

    • BlogPosts
      • Audio Features
      • Classifier Metrics
      • Course Work
      • Deep Learning Optimizers
      • Environment Setup
      • KMeans Clustering
      • Resources
      • Understanding Latents: Variational Auto-Encoder
    • GradSchool
      • GradSchool
    • Research Wikis
      • Audio LLM Wiki
        • Concepts
          • Audio-Text Pretraining Patterns
          • Encoder-Free Early Fusion
          • Full-Duplex Spoken Dialogue
          • Inner Monologue
          • Interaction Models
          • Interaction-Background Model Split
          • Multi-Stream Audio Modeling
          • RQ-Transformer
          • Speech Understanding
          • Thinker-Talker Architecture
          • Time-Aligned Micro-Turns
        • Entities
          • AuT (Audio Transformer)
          • FD-bench
          • Helium
          • Mimi (neural audio codec)
          • Moshi
          • Qwen3-Omni
          • TML-Interaction-Small
          • Voxtral
        • Sources
          • Interaction Models: A Scalable Approach to Human-AI Collaboration
          • Moshi: a speech-text foundation model for real-time dialogue
          • Qwen3-Omni (technical report)
          • Voxtral
        • Topics
          • Real-Time Interactive Speech Models
    Home

    ❯

    Research Wikis

    ❯

    Audio LLM Wiki

    ❯

    Entities

    Folder: wiki/mm-llm-wiki/Entities

    8 items under this folder.

    • Jul 22, 2026

      AuT (Audio Transformer)

      • entity
      • model
      • audio
      • speech
      • audio-encoder
    • Jul 22, 2026

      FD-bench

      • entity
      • benchmark
      • interactivity
    • Jul 22, 2026

      Helium

      • entity
      • model
      • text-llm
    • Jul 22, 2026

      Mimi (neural audio codec)

      • entity
      • model
      • audio-codec
      • speech
    • Jul 22, 2026

      Moshi

      • entity
      • model
      • speech
      • full-duplex
    • Jul 22, 2026

      Qwen3-Omni

      • entity
      • model
      • multimodal
      • omni-modal
      • audio
      • speech
      • full-duplex
      • open-weights
    • Jul 22, 2026

      TML-Interaction-Small

      • entity
      • model
      • multimodal
    • Jul 22, 2026

      Voxtral

      • entity
      • model
      • speech
      • audio-understanding
      • multimodal
      • open-weights

    Created with Quartz v4.4.0 © 2026

    • GitHub
    • Discord Community