fix(vad): clamp cursor when resizing speech buffer - #7566
Open
gudcks0305 wants to merge 2 commits into
Open
gudcks0305 wants to merge 2 commits into
gudcks0305 wants to merge 2 commits into
Conversation
Contributor
There was a problem hiding this comment.
🔍 Devin Review: 1 flag
Not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Reducing
max_buffered_speechorprefix_padding_durationwhile a VAD stream has buffered speech can shrink its sample array below its write cursor. The next speech event then raisesValueError: data length must be >= num_channels * samples_per_channel * sizeof(int16)and terminates the stream. Shrinking and immediately growing the buffer also leaves discarded samples counted as valid audio.Change
Keep the cursor on the stream and clamp it when resizing the buffer, in both
inference.VADand the Silero plugin. Existing event emission, prefix retention, and reset paths use the same cursor. Add deterministic regressions for both options, buffer growth, shrink-then-grow, subsequent speech, and explicit reset.Load the optional Silero plugin only for its fixture cases, so its absence does not prevent the bundled VAD regressions from running. Skip missing modules while allowing a general
ImportErrorfrom a broken installed plugin to propagate.Validation
rtc.AudioFramebuffer paths.ImportErrorpropagated as 6 fixture errors with a nonzero exit status, while the 6 bundled cases still passed.The local environment installed core Agents plus the Silero extra. Full
--unit --audio_eotcollection stopped at missing optional provider packages, and the project-wide type check was blocked by missing optional packages and typing stubs. Focused mypy checked the two changed production modules with--follow-imports=silent. No cloud/provider integration tests were run.AI assistance was used to develop and test this change.