I speed up the generation part of the demo in case you get bored 😄
I also tested another long-form generation, and the VRAM usage looks stable. The demo is about a minute long, and I posted it on X.
This started as a random idea and somehow turned into a full detour from working on the next audio.cpp release. The model was uploaded to the audio.cpp HF repo. I will upload the xcframework later, and then push the code to a branch after release 0.6.
Link: https://v.redd.it/23bwnqd2sghh1
Subreddit: r/LocalLLaMA