Kokoro 82M β Backpack Voice Package
π Backpack Verified
Lightweight local text-to-speech with a curated set of voices. This package stages immutable upstream artifacts for Backpack's voice runtime layer. It does not replace the chat model selected by the user.
Package
| Field | Value |
|---|---|
| Capability | text-to-speech |
| Input | text |
| Output | audio |
| Runtime | kokoro |
| Configured runtime revision | kokoro==0.9.4;en_core_web_sm==3.8.0 |
| Format | pytorch | | Precision | F32 | | Package size | 314.1 MiB | | Recommended RAM | 1.88 GB | | Languages | multilingual |
Validation status
The packager verified the immutable revision, selected-file inventory, non-empty files, hashes, configuration JSON, and the primary artifact container/header. It also loaded the package and passed deterministic audio inference with the configured runtime.
| Integrity | Metadata | Runtime load | Audio inference | Tokenizer |
|---|---|---|---|---|
| passed | passed | passed | passed | passed |
Run with Kokoro
Install kokoro==0.9.4, construct KModel from config.json and kokoro-v1_0.pth,
then pass one of the packaged voices/*.pt files to KPipeline.
Provenance
Upstream: hexgrad/Kokoro-82M
Immutable revision:
f3ff3571791e39611d31c381e3a41a3af07b4987License:
apache-2.0Backpack copied the selected upstream artifacts without modifying model weights.
Backpack did not train this model and does not claim ownership of it.
Files and checksums
config.jsonβ 2.3 KiB β5abb01e2403b072bf03d04fde160443e209d7a0dad49a423be15196b9b43c17fkokoro-v1_0.pthβ 312.1 MiB β496dba118d1a58f5f3db2efc88dbdc216e0483fc89fe6e47ee1f2c53f18ad1e4voices/af_heart.ptβ 511.2 KiB β0ab5709b8ffab19bfd849cd11d98f75b60af7733253ad0d67b12382a102cb4ffvoices/am_michael.ptβ 511.2 KiB β9a443b79a4b22489a5b0ab7c651a0bcd1a30bef675c28333f06971abbd47bd37voices/bf_emma.ptβ 511.2 KiB βd0a423deabf4a52b4f49318c51742c54e21bb89bbbe9a12141e7758ddb5da701voices/bm_george.ptβ 511.2 KiB βf1bc812213dc59774769e5c80004b13eeb79bd78130b11b2d7f934542dab811b
Review the upstream model card and license before use or redistribution. Speech systems can mis-transcribe, synthesize misleading content, or behave differently across languages and accents.
- Downloads last month
- -