Activity Feed

AI & ML interests

Pretraining, hybrid architectures, exploring exotic training methods at small scale.

Recent Activity

arthurblg1802  updated a Space about 2 hours ago
VantoraLabs/README
arthurblg1802  updated a model about 3 hours ago
VantoraLabs/Vocetta2-1m
arthurblg1802  published a model about 3 hours ago
VantoraLabs/Vocetta2-1m
View all activity

Organization Card

Vantora Labs

Vantora Labs builds small, measured models and publishes everything: weights, runtime code, training and benchmark scripts, and the honest numbers they produce. Our flagship line is Vocetta2, a state-of-the-art English text-to-speech family, and the smallest and fastest SOTA TTS system we know of.

Vocetta2-1m — a complete TTS pipeline in 1,031,754 parameters, running ~90× real time on CPU (RTF 0.011), scoring SCOREQ 4.04, with no GPU required.

Vocetta2 — SOTA TTS from 147K to 1M parameters

A full four-stage speech stack: dictionary-first grapheme-to-phoneme front end, duration predictor, acoustic model, waveform decoder. Distilled stage by stage from a larger teacher. Every member of the family is MIT-licensed, runs faster than real time on CPU, and needs nothing but Python and three pip packages.

Model Parameters Speed (CPU) WER ↓ SCOREQ ↑
Vocetta2-1m 1,031,754 ~90× RT (RTF 0.011) 0.083 4.04
Vocetta2-500k 499,528 ~125× RT (RTF 0.008) 0.117 2.70
Vocetta2-276k 276,765 ~80× RT (RTF 0.0125) 0.101 2.04
Vocetta2-147k 147,532 ~134× RT (RTF 0.0075) 0.105 1.48

WER is word error rate under Whisper-small transcription on the 24-sentence diverse held-out set; SCOREQ rates naturalness. The first-generation Vocetta-181K opened the line at 181,189 parameters; across the two generations, diverse-set WER improved from 0.23 to 0.083 and SCOREQ from 1.03 to 4.04.

Start with the 1M flagship: its samples/ directory holds eight clips rendered by the exact shipped checkpoint, and the card documents the full architecture and every measurement behind the table.

Vantora Micro — architectures and computational cost

Before the Vocetta line, the Vantora-Micro models probed architectures and computational cost at a fixed training budget: a pure Llama-style transformer against a Mamba-2 + attention hybrid, both trained on the same 100M-token FineWeb-Edu slice with the same batch size and step count.

Model Task Parameters Result
Vantora-Micro Text generation 9,800 BananaMind Elo 810, 26.0% accuracy
Vantora-Micro-Hybrid Text generation 11,256 BananaMind Elo 863, 30.3% accuracy

The hybrid won 6 of 7 benchmark categories but trained 11.5× slower, and on PIQA, HellaSwag and ARC-Easy the difference sits within the noise floor at this scale. The documented conclusion is that the SSM does not justify its training cost here, though it was a clear win on TinyStories. Both model cards carry the full methodology and limitations.

Computing

All training and evaluation runs on a single workstation: an NVIDIA GTX 750 (Maxwell, 4 GB VRAM, no tensor cores), a Ryzen 5 3500X, and 16 GB of RAM. The constraint is deliberate: at these sizes a full training run takes minutes to hours, so architecture and optimization decisions can be tested rather than assumed.

Reporting

Speech numbers use Whisper small as the transcription judge for word error rate, and SCOREQ and DNSMOS for naturalness; language model scores use the official BananaMind runner on its verified split. Each result is reported against the baseline it should be compared with, and limitations are documented on the model cards themselves.

Contact

Issues and questions are welcome on any model repository.

datasets 0

None public yet