🪴 Prabhav's Log

    • BlogPosts
      • Audio Features
      • Classifier Metrics
      • Course Work
      • Deep Learning Optimizers
      • Environment Setup
      • KMeans Clustering
      • Resources
      • Understanding Latents: Variational Auto-Encoder
    • GradSchool
      • GradSchool
    • Research Wikis
      • Audio LLM Wiki
        • Concepts
          • Audio-Text Pretraining Patterns
          • Encoder-Free Early Fusion
          • Full-Duplex Spoken Dialogue
          • Inner Monologue
          • Interaction Models
          • Interaction-Background Model Split
          • Multi-Stream Audio Modeling
          • RQ-Transformer
          • Speech Understanding
          • Thinker-Talker Architecture
          • Time-Aligned Micro-Turns
        • Entities
          • AuT (Audio Transformer)
          • FD-bench
          • Helium
          • Mimi (neural audio codec)
          • Moshi
          • Qwen3-Omni
          • TML-Interaction-Small
          • Voxtral
        • Sources
          • Interaction Models: A Scalable Approach to Human-AI Collaboration
          • Moshi: a speech-text foundation model for real-time dialogue
          • Qwen3-Omni (technical report)
          • Voxtral
        • Topics
          • Real-Time Interactive Speech Models
    Home

    ❯

    tags

    ❯

    Tag: multimodal

    Tag: multimodal

    11 items with this tag.

    • Jul 22, 2026

      Speech Understanding

      • concept
      • speech
      • multimodal
      • architecture
    • Jul 22, 2026

      Qwen3-Omni

      • entity
      • model
      • multimodal
      • omni-modal
      • audio
      • speech
      • full-duplex
      • open-weights
    • Jul 22, 2026

      TML-Interaction-Small

      • entity
      • model
      • multimodal
    • Jul 22, 2026

      Voxtral

      • entity
      • model
      • speech
      • audio-understanding
      • multimodal
      • open-weights
    • Jul 22, 2026

      Interaction Models: A Scalable Approach to Human-AI Collaboration

      • source
      • interaction-models
      • multimodal
      • real-time
    • Jul 22, 2026

      Moshi: a speech-text foundation model for real-time dialogue

      • source
      • speech
      • full-duplex
      • real-time
      • multimodal
    • Jul 22, 2026

      Qwen3-Omni (technical report)

      • source
      • paper
      • multimodal
      • audio
      • speech
    • Jul 22, 2026

      Voxtral

      • source
      • speech
      • audio-understanding
      • multimodal
      • open-weights
    • Jul 22, 2026

      Audio-Text Pretraining Patterns

      • concept
      • training
      • speech
      • multimodal
    • Jul 22, 2026

      Encoder-Free Early Fusion

      • concept
      • architecture
      • multimodal
    • Jul 22, 2026

      Interaction Models

      • concept
      • real-time
      • multimodal
      • human-ai-collaboration

    Created with Quartz v4.4.0 © 2026

    • GitHub
    • Discord Community