Inference Optimization
Topic archive • 2 matches
2026-09-07
Technology
New inference-time method reduces hallucinations in OpenAI's Whisper: Researchers have developed a training-free, inference-time method to reduce hallucinated transcripts in OpenAI's Whisper model. The approach estimates a compact hallucination-associated subspace from non-speech calibration data and projects decoder hidden states away from it.
AI Research • arXiv
Permalink
2026-09-04
Technology
SMC mechanism speeds up tool-using LLM agents via speculative drafting: Researchers have introduced Speculative Macro Commit (SMC), a runtime mechanism designed to reduce wall-clock time for tool-using LLM agents. SMC uses a faster speculative drafter model to predict and execute future action chains on an isolated environment snapshot.
Research • arXiv
Permalink