From a31746dea0d5e5df07ddc1026b08ac9fd03c1b53 Mon Sep 17 00:00:00 2001 From: LauraGPT <18321252+LauraGPT@users.noreply.github.com> Date: Wed, 30 Sep 2026 16:36:43 +0000 Subject: [PATCH] docs: surface merged speech-to-speech SenseVoice integration Signed-off-by: LauraGPT <18321252+LauraGPT@users.noreply.github.com> --- README.md | 2 +- README_zh.md | 2 +- docs/community_growth_20k.md | 2 +- docs/community_projects.md | 21 +++++++++++++ docs/community_projects_zh.md | 21 +++++++++++++ tests/test_community_projects_entries.py | 21 +++++++++++++ .../product-site/content/legacy-manifest.json | 4 +-- web-pages/product-site/legacy/ecosystem.html | 5 +++ .../product-site/legacy/en/ecosystem.html | 5 +++ .../product-site/tests/test_documentation.py | 31 +++++++++++++++++++ 10 files changed, 109 insertions(+), 5 deletions(-) diff --git a/README.md b/README.md index a5ca0f6e5b..df4c3cfde9 100644 --- a/README.md +++ b/README.md @@ -125,7 +125,7 @@ results = model.generate(["audio1.wav", "audio2.wav"], language="auto") > > **Use with AI agents:** [MCP Server](examples/mcp_server/) for Claude/Cursor · [OpenAI API](examples/openai_api/) for LangChain/Dify/AutoGen > -> **Use with voice agents:** [OpenClaw realtime plugin](integrations/openclaw/) for self-hosted Talk and Voice Call transcription +> **Use with voice agents:** [OpenClaw realtime plugin](integrations/openclaw/) for self-hosted Talk and Voice Call transcription · [Hugging Face speech-to-speech with SenseVoice](./docs/community_projects.md#speech-to-speech) (merged source checkout; not in v1.0.0) ### Why FunASR? diff --git a/README_zh.md b/README_zh.md index d6e9bb7871..5d6e439159 100644 --- a/README_zh.md +++ b/README_zh.md @@ -115,7 +115,7 @@ results = model.generate(["audio1.wav", "audio2.wav"], language="auto") > > **接入 AI Agent:** [MCP 服务](examples/mcp_server/) 支持 Claude/Cursor · [OpenAI API](examples/openai_api/README_zh.md) 支持 LangChain/Dify/AutoGen > -> **接入语音 Agent:** [OpenClaw 实时转写插件](integrations/openclaw/) 支持私有部署的 Talk 与 Voice Call 转写 +> **接入语音 Agent:** [OpenClaw 实时转写插件](integrations/openclaw/) 支持私有部署的 Talk 与 Voice Call 转写 · [Hugging Face speech-to-speech + SenseVoice](./docs/community_projects_zh.md#speech-to-speech)(已合入源码,v1.0.0 尚未包含) ### 为什么选 FunASR? diff --git a/docs/community_growth_20k.md b/docs/community_growth_20k.md index ca210fb00b..9c72d1b4db 100644 --- a/docs/community_growth_20k.md +++ b/docs/community_growth_20k.md @@ -293,7 +293,7 @@ High-star feature requests and roadmap issues are earlier in the funnel than PRs | `ray-project/ray#64053` Ray Serve FunASR ASR example | Puts FunASR in production serving docs for teams already using Ray | Current head `9b2a3b342e85` is mergeable but still has two red external gates. `docs/readthedocs.com:anyscale-ray` fails on unrelated Tune collection references outside the FunASR files. `buildkite/microcheck` includes a repository-wide dashboard ESLint/environment failure plus one PR-local Black issue: `doc/source/serve/doc_code/funasr_asr.py` would be reformatted. The one-line formatting fix is already isolated in `nh-atuan/ray#3`, which is open, clean, and mergeable at `559820895d47`; local validation passes `py_compile`, the two FunASR doc-code tests, `black --check`, and `git diff --check`. A direct 2026-07-23 maintainer push to `nh-atuan:issue-64052` was rejected with `403 Permission to nh-atuan/ray.git denied to LauraGPT`, so the next action is for the contributor or a Ray maintainer to merge `nh-atuan/ray#3` before rerunning the Ray PR checks. | | `huggingface/optimum-intel#1801` OpenVINO support | Helps CPU and edge users evaluate Fun-ASR on Intel hardware | Merged after the FunASR-side review and validation pass. Treat this as a completed Intel/OpenVINO discovery and runtime win; watch downstream OpenVINO releases and user issues rather than keeping the closed PR in the default active queue. | | `huggingface/optimum-intel#1874` FunASR OpenVINO export tests | Makes FunASR export coverage visible in Optimum Intel's parameterized precommit suite rather than relying on an unverified model mapping | PR is approved, mergeable, and green. Comparing base `109314ca` with head `1dafe5e1` shows three new FunASR cases collected: one export case plus compressed- and uncompressed-weight quantization cases. The corresponding remote precommit jobs selected and passed those cases; exact collection and log evidence is at https://github.com/huggingface/optimum-intel/pull/1874#issuecomment-5054273045. Wait for maintainer merge. | -| `huggingface/speech-to-speech#319` SenseVoice STT handler | Adds SenseVoice/FunASR to local open-source voice-agent pipelines where low-latency STT is a core comparison point | Head `3657a9bc2782` is mergeable with no unresolved review threads. Ruff, formatting, compilation, and the focused CLI/Paraformer/SenseVoice/STT suite pass (`22 passed`); the real `setup -> warmup -> process` path also transcribed a six-second clip on H100 and preserved the final event metadata. Wait for maintainer review without another status bump. | +| `huggingface/speech-to-speech#319` SenseVoice STT handler | Makes native SenseVoice/FunASR available in Hugging Face's voice-agent pipeline | Merged by the upstream maintainer on 2026-09-29 (UTC) as `9e2ed1099190a4e4bc8a972b4a3488949ff1b9f6`, with approval on head `4e02e0a5511db7f0ba2c3850786db6dea6a9eee8`. As checked 2026-09-30, released `v1.0.0` does not contain the backend: use the [pinned source recipe](./community_projects.md#speech-to-speech). Remove the review-wait action; track a release containing the integration instead. Historical model/hardware tests are not new validation of this merged revision. | | `OpenBMB/VoxCPM#349` Windows CUDA installer with SenseVoice fallback | Puts SenseVoice into a 32k-star multilingual TTS app's Windows install and first-run ASR path as the fallback when local Parakeet/CUDA is unavailable | Current head `9f34141cc81e04028c4e62ac652f2a66dd453dfa` is dirty only in `app.py`; LauraGPT posted the conflict recipe and light validation at https://github.com/OpenBMB/VoxCPM/pull/349#issuecomment-4905293176. Wait for author rebase or maintainer action without duplicate comments. | | `livekit/agents#6176` FunASR/SenseVoice realtime STT plugin | Opens a path into LiveKit's realtime voice-agent ecosystem where local STT is evaluated alongside hosted providers | Current head `9f995f7ee5fc` is mergeable and review-gated with all visible checks green or skipped. Fresh 2026-07-23 validation passed the 3 focused FunASR plugin tests, Ruff, format check, `uv lock --check`, `compileall`, and `git diff --check`; evidence was posted at https://github.com/livekit/agents/pull/6176#issuecomment-5051589686. Avoid duplicate pings and monitor for maintainer review on plugin scope, package metadata, or optional dependency expectations. | | `datajuicer/data-juicer#938` HumanVBench audio/video operators | Places FunASR/SenseVoice-style speech understanding into a 6k-star data processing toolkit used to evaluate human-centric video and multimodal datasets | Current PR is open and mergeable but process-blocked while unit-test jobs are still waiting; LauraGPT posted the FunASR/SenseVoice dependency and validation follow-up at https://github.com/datajuicer/data-juicer/pull/938#issuecomment-4905235851. Wait for CI or maintainer feedback before adding more comments. | diff --git a/docs/community_projects.md b/docs/community_projects.md index 168cbdc381..1777f61684 100644 --- a/docs/community_projects.md +++ b/docs/community_projects.md @@ -6,6 +6,27 @@ This page lists maintained projects where FunASR, Fun-ASR-Nano, SenseVoice, or a These integrations are community-maintained. Their release cadence, hardware support, and API stability are controlled by the upstream project, not by the FunASR maintainers. + +## Hugging Face speech-to-speech: SenseVoice + +The native SenseVoiceSmall backend was merged upstream in [#319](https://github.com/huggingface/speech-to-speech/pull/319) on 2026-09-29 (UTC). It transcribes VAD segments through FunASR, uses the Hugging Face checkpoint `FunAudioLLM/SenseVoiceSmall`, and emits the pipeline's progressive/final transcription events. These events do not make the offline checkpoint a chunk-streaming ASR model. The checkpoint supports Mandarin, Cantonese, English, Japanese, and Korean. + +**Source installation, checked 2026-09-30:** this backend is not included in the released [v1.0.0](https://github.com/huggingface/speech-to-speech/releases/tag/v1.0.0). Use the merged source revision below, not `pip install speech-to-speech==1.0.0`. First follow the upstream [platform and environment setup](https://github.com/huggingface/speech-to-speech/blob/9e2ed1099190a4e4bc8a972b4a3488949ff1b9f6/README.md); Python 3.10+ is required. + +```bash +git clone https://github.com/huggingface/speech-to-speech.git +cd speech-to-speech +git checkout --detach 9e2ed1099190a4e4bc8a972b4a3488949ff1b9f6 +python -m pip install -e ".[sensevoice]" +speech-to-speech serve --stt sense-voice \ + --sense_voice_stt_device cpu \ + --sense_voice_stt_model_name FunAudioLLM/SenseVoiceSmall --help +``` + +The last command inspects CLI options; it does not start a model-backed conversation. After configuring the LLM and TTS backends using the upstream guide, add these SenseVoice flags (without `--help`) to your `serve` or `local` command. The explicit CPU flag avoids assuming a compatible CUDA installation; model files are downloaded on first actual use unless cached. + +Local STT does not imply an entirely offline voice agent: LLM/TTS backends and their network or API-key requirements are configured separately. The upstream `--log_transcripts` option controls application logs, not the terminal conversation display; terminal output may still contain transcription text. See the pinned [handler](https://github.com/huggingface/speech-to-speech/blob/9e2ed1099190a4e4bc8a972b4a3488949ff1b9f6/src/speech_to_speech/STT/sense_voice_handler.py) and [logging contract](https://github.com/huggingface/speech-to-speech/blob/9e2ed1099190a4e4bc8a972b4a3488949ff1b9f6/src/speech_to_speech/pipeline/transcript_logging.py). No new model, GPU, or complete voice-conversation validation is claimed by this listing. + ## Voice agents and applications | Project | What is integrated | Start here | diff --git a/docs/community_projects_zh.md b/docs/community_projects_zh.md index c23927a30d..e2f72eadf6 100644 --- a/docs/community_projects_zh.md +++ b/docs/community_projects_zh.md @@ -6,6 +6,27 @@ 这些集成由社区项目维护,其发布节奏、硬件支持和 API 稳定性由上游项目决定,不属于 FunASR 维护者的兼容性承诺。 + +## Hugging Face speech-to-speech:SenseVoice + +原生 SenseVoiceSmall 后端已于 2026-09-29(UTC)通过 [#319](https://github.com/huggingface/speech-to-speech/pull/319) 合入上游。它通过 FunASR 转写 VAD 分段,使用 Hugging Face checkpoint `FunAudioLLM/SenseVoiceSmall`,并输出流水线的 progressive/final 转写事件。这些事件不代表离线 checkpoint 具备分块流式 ASR 能力。该 checkpoint 支持普通话、粤语、英语、日语和韩语。 + +**源码安装,核验日期 2026-09-30:** 已发布的 [v1.0.0](https://github.com/huggingface/speech-to-speech/releases/tag/v1.0.0) 尚未包含此后端。请使用下面已合入的源码提交,不要用 `pip install speech-to-speech==1.0.0` 代替。先按上游的[平台与环境配置](https://github.com/huggingface/speech-to-speech/blob/9e2ed1099190a4e4bc8a972b4a3488949ff1b9f6/README.md)准备环境;需要 Python 3.10+。 + +```bash +git clone https://github.com/huggingface/speech-to-speech.git +cd speech-to-speech +git checkout --detach 9e2ed1099190a4e4bc8a972b4a3488949ff1b9f6 +python -m pip install -e ".[sensevoice]" +speech-to-speech serve --stt sense-voice \ + --sense_voice_stt_device cpu \ + --sense_voice_stt_model_name FunAudioLLM/SenseVoiceSmall --help +``` + +最后一条命令只查看 CLI 选项,不启动真实模型对话。按上游指南配置 LLM 与 TTS 后,把这些 SenseVoice 参数(去掉 `--help`)加入自己的 `serve` 或 `local` 命令。这里显式选择 CPU,不假定已安装兼容 CUDA 环境;首次实际运行会下载未缓存的模型文件。 + +本地 STT 不等于整条语音 Agent 流水线都离线:LLM/TTS 后端及网络、API key 要求需要单独配置。上游 `--log_transcripts` 控制应用日志,不控制终端对话显示;终端输出仍可能包含转写文本。实现边界见固定版本的 [handler](https://github.com/huggingface/speech-to-speech/blob/9e2ed1099190a4e4bc8a972b4a3488949ff1b9f6/src/speech_to_speech/STT/sense_voice_handler.py) 和[日志契约](https://github.com/huggingface/speech-to-speech/blob/9e2ed1099190a4e4bc8a972b4a3488949ff1b9f6/src/speech_to_speech/pipeline/transcript_logging.py)。本条目不宣称完成了新的模型、GPU 或完整语音对话验证。 + ## 语音 Agent 与应用 | 项目 | 已集成能力 | 从这里开始 | diff --git a/tests/test_community_projects_entries.py b/tests/test_community_projects_entries.py index b9f96fb544..28c76a8c65 100644 --- a/tests/test_community_projects_entries.py +++ b/tests/test_community_projects_entries.py @@ -32,3 +32,24 @@ def test_recent_merged_discovery_lists_are_listed(): for text in [english, chinese]: assert "WangRongsheng/awesome-LLM-resources" in text assert "WangRongsheng/awesome-LLM-resources/pull/162" in text + + +def test_speech_to_speech_source_install_is_discoverable_and_release_bounded(): + revision = "9e2ed1099190a4e4bc8a972b4a3488949ff1b9f6" + for suffix in ("", "_zh"): + text = (ROOT / f"docs/community_projects{suffix}.md").read_text() + assert 'id="speech-to-speech"' in text + assert "https://github.com/huggingface/speech-to-speech/pull/319" in text + assert f"git checkout --detach {revision}" in text + assert 'python -m pip install -e ".[sensevoice]"' in text + assert "speech-to-speech serve --stt sense-voice" in text + assert "--sense_voice_stt_device cpu" in text + assert "--sense_voice_stt_model_name FunAudioLLM/SenseVoiceSmall" in text + assert "--help" in text + assert "v1.0.0" in text and "2026-09-30" in text + readme = (ROOT / f"README{suffix}.md").read_text() + assert f"./docs/community_projects{suffix}.md#speech-to-speech" in readme + row = next(line for line in (ROOT / "docs/community_growth_20k.md").read_text().splitlines() + if "`huggingface/speech-to-speech#319`" in line) + assert revision in row and "Merged" in row + assert "Wait for maintainer review" not in row diff --git a/web-pages/product-site/content/legacy-manifest.json b/web-pages/product-site/content/legacy-manifest.json index c3c5c23736..ddfe0410d0 100644 --- a/web-pages/product-site/content/legacy-manifest.json +++ b/web-pages/product-site/content/legacy-manifest.json @@ -53,7 +53,7 @@ "css/index.1682178c.css": "b5f6034ce886c25773928c114e4290d1c491a3d081677b488810198efd76b4e1", "decoder.js": "f2023a28036b1ff58e5bdabbb72ea427f8faaf60e89acd78ff183c313bdb0a2c", "donors.html": "eec6b789b29c27d372925ae0ad3e13d4c6e62978b3c38c4da9e3aec129f10934", - "ecosystem.html": "a3cdaf0fefdc10fd51e8f09c3239e338528f4700e0037c2eb9556d650f6faef9", + "ecosystem.html": "27ad59792848d3060c52afb3d794ae1d98c553e93ba99932c7101f4eae09639c", "en/blog/cantonese-speech-recognition.html": "b9bd146198206f781c37f341c370a690dc83cd7163334e34b4e7f65e69af8ff2", "en/blog/chinese-speech-recognition.html": "16b68bd2a1122236c580c62275002ae2c0875ca14e32686e88d3b5ccfa9f56ef", "en/blog/fun-asr-nano-guide.html": "1746c708bd0606188de4266d8255984dd9e14634af33c54ad65a5f05b767be1b", @@ -91,7 +91,7 @@ "en/blog/voice-activity-detection-python.html": "d12f5d0f4ae07c3434f6188d12f5603b8abb9578413b79133668638d1c5e18a7", "en/blog/which-funasr-model.html": "1207180aa929a3aa51db136c1098f4edc8f07405e2b56fe033998bf996f8ec7f", "en/donors.html": "eacb7f9434b1b29e13e9b0175bb9680556ee72ab4dbd155e68f9c912eb0212b1", - "en/ecosystem.html": "38356f6bdb5cd89988304be61f13707f4df9fa84c2a191701aaee475d7587902", + "en/ecosystem.html": "fae2a46e7fbbb702c36d06e2f6cf0bf3777986e7333ea6684a12830cb392158d", "en/index.html": "d86effc3bf4d218eabd465cbbe026f62c98dd50110ddd58251d9d271a64882a2", "en/llama-cpp.html": "dbfbd125ebbf07f7d923c91b122818f8b09484ce54c1e80eba66a89d22dfb5b4", "en/models.html": "9806d01ba1807c2a6a88f7232102e185426fb8458bf51e1d7a8ce872352e1fd1", diff --git a/web-pages/product-site/legacy/ecosystem.html b/web-pages/product-site/legacy/ecosystem.html index 6b30b9be9a..5b3625219f 100644 --- a/web-pages/product-site/legacy/ecosystem.html +++ b/web-pages/product-site/legacy/ecosystem.html @@ -280,6 +280,11 @@