re-clarify the opensource version weight
Browse files
README.md
CHANGED
|
@@ -21,6 +21,8 @@ Kangdi Wang<sup>1</sup> · Yusheng Dai<sup>2</sup> · Jin Xu<sup>1†</sup>
|
|
| 21 |
|
| 22 |
[[Demo Page](https://eps-acoustic-revolution-lab.github.io/EAR_VAE2/)] - [[Paper](https://arxiv.org/abs/2608.19843)] - [[Codebase](https://github.com/Eps-Acoustic-Revolution-Lab/EAR_VAE2)]
|
| 23 |
|
|
|
|
|
|
|
| 24 |
---
|
| 25 |
|
| 26 |
<p align="center">
|
|
@@ -126,21 +128,6 @@ python inference.py \
|
|
| 126 |
--config checkpoints/config.json \
|
| 127 |
--input input.wav --output output.wav
|
| 128 |
```
|
| 129 |
-
|
| 130 |
-
|
| 131 |
-
## Model Details
|
| 132 |
-
|
| 133 |
-
| Config | Params (M) | Latent dim | Rate (Hz) | Compression |
|
| 134 |
-
|--------|:---:|:---:|:---:|:---:|
|
| 135 |
-
| Small (C0=64) | ~42.6 | 128 | 25 | 1920× |
|
| 136 |
-
|
| 137 |
-
- **Sample rate**: 48 kHz stereo
|
| 138 |
-
- **STFT**: 3840-point FFT, 1920-sample hop → 25 Hz frame rate
|
| 139 |
-
- **Latent**: 128-d continuous (VAE with KL regularization)
|
| 140 |
-
- **Refiner**: 12-layer banded Transformer (256-d, 1024 intermediate)
|
| 141 |
-
|
| 142 |
-
> **⚠️ Note on open-source weights:** Due to data licensing constraints, the open-source model weights are **retrained on publicly available datasets** (not the full internal training corpus). Performance may differ from the numbers reported in the paper, which were obtained with the full-scale proprietary training data.
|
| 143 |
-
|
| 144 |
---
|
| 145 |
|
| 146 |
## Citation
|
|
|
|
| 21 |
|
| 22 |
[[Demo Page](https://eps-acoustic-revolution-lab.github.io/EAR_VAE2/)] - [[Paper](https://arxiv.org/abs/2608.19843)] - [[Codebase](https://github.com/Eps-Acoustic-Revolution-Lab/EAR_VAE2)]
|
| 23 |
|
| 24 |
+
> **⚠️ Note on open-source weights:** Due to data licensing constraints, the open-source model weights are **retrained on publicly available datasets** (not the full internal training corpus). Performance may differ from the numbers reported in the paper, which were obtained with the full-scale proprietary training data.
|
| 25 |
+
|
| 26 |
---
|
| 27 |
|
| 28 |
<p align="center">
|
|
|
|
| 128 |
--config checkpoints/config.json \
|
| 129 |
--input input.wav --output output.wav
|
| 130 |
```
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 131 |
---
|
| 132 |
|
| 133 |
## Citation
|