I discovered a research paper, I'll try to implement what they demonstrated, and then experiment to make it use as less resource as possible
> DuplexCascade: Full-Duplex Speech-to-Speech Dialogue with VAD-Free Cascaded ASR–LLM–TTS Pipeline and Micro-Turn Optimization
> https://arxiv.org/html/2603.09180
Would be fun to have a full duplex with a local LLM :D