Skip to content

Feature: 集成 SenseVoice/FunASR 离线语音转文字 #286

Description

@LauraGPT

Important

更正: 下文早期的横向速度和准确率说法没有来自同配置基准,现予以撤回。SenseVoiceSmall 支持中文、粤语、英语、日语和韩语,可返回语言、情感和音频事件标签;说话人分离需要另接 CAM++ 等独立模型或现有 pipeline,并非 SenseVoice 的内置输出。运行速度、时间戳、标点和效果取决于具体模型、接口、硬件和音频。FunASR 与 SenseVoice 仓库源码采用 MIT,模型权重以各自模型卡为准。是否集成应以本项目实际数据测试为准。

你好!feishu-openai 是一个很棒的项目。

建议考虑集成 SenseVoice / FunASR 作为语音转文字后端:

SenseVoice 的优势:

  • 比 Whisper 快约 10 倍(GPU 上 10 秒音频仅需 50ms)
  • 50+ 语言支持:中英日韩等
  • 完全本地部署:无需调用 OpenAI API,数据安全
  • 内置 VAD + 标点恢复
  • OpenAI 兼容 API:
pip install funasr
funasr-server --device cuda
# 与 OpenAI /v1/audio/transcriptions 兼容
curl http://localhost:8000/v1/audio/transcriptions -F file=@audio.wav

对于飞书场景,用户发送语音消息后可以用 FunASR 快速转写为文字,再交给 LLM 处理。完全本地化部署也更适合企业内部使用。

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions