SyncNet and S3FD weights

This repository mirrors two pretrained weight files from the Oxford VGG SyncNet release, so that tools can download them from a stable URL. The files are byte-identical to the ones linked from joonson/syncnet_python.

File Model Size SHA-256
syncnet_v2.model SyncNet audio-visual synchronization network (Chung and Zisserman, 2016) 54.6 MB 961e8696f888fce4f3f3a6c3d5b3267cf5b343100b238e79b2659bff2c605442
sfd_face.pth S3FD face detector, as distributed with the SyncNet code 89.8 MB d54a87c2b7543b64729c9a25eafd188da15fd3f6e02f0ecec76ae1b30d86c491

Used by

  • syncnet-python (pip install syncnet-python): SyncNet for Python 3.9 to 3.13 and PyTorch 2. It returns the confidence and distance values reported as LSE-C and LSE-D in lip-sync and talking-head papers.
  • TalkNet-ASD-py312: TalkNet active speaker detection for Python 3.12 and PyTorch 2. It downloads sfd_face.pth from this repository and checks the SHA-256 above before loading it.

Terms

The Oxford VGG page states that the SyncNet model "can be used for research purposes under Creative Commons Attribution License". The original authors do not state separate terms for sfd_face.pth. Please check the original release before using either file outside research.

Citation

@InProceedings{Chung16a,
  author       = "Chung, J.~S. and Zisserman, A.",
  title        = "Out of time: automated lip sync in the wild",
  booktitle    = "Workshop on Multi-view Lip-reading, ACCV",
  year         = "2016",
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support