YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
pCoMole: Pareto-Constrained Molecule Editing with Discrete Flows
Repository layout
pCoMole/
train.py / evaluate.py # Edit Flow train / val-loss eval
configs/ # GFP (protein) and SELFIES configs
model/ logic/ flow_matching/ # Edit Flow implementation
smiles_tokenizer/ # SMILES SPE + SELFIES vocab
data/selfies/28k_mimetics/ # shipped SELFIES training split
gfp/ # GFP generation + pCoMole
pcomole.py # official GFP editor
generate.py
ckpt/last_2.ckpt
classifier_ckpt/best.pt
FPredX/ # excitation / brightness / emission models
peptidomimetics/ # peptidomimetic generation + pCoMole
pcomole.py
generate.py
objectives.py # PeptiVerse + Admetica/DeepDTAGen mix
ckpt/SELFIES_EditFlows.ckpt
ckpt/SMILES_BindEvaluator.ckpt
ckpt/admetica/ # LD50, solubility, Caco-2, half-life
ckpt/deepdtagen/
1. Environment
A CUDA GPU is strongly recommended.
git clone <this-repo-url> pCoMole
cd pCoMole
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
HuggingFace model downloads happen on first use (facebook/esm2_t33_650M_UR50D, aaronfeller/PeptideCLM-23M-all, ChemBERTa).
PeptiVerse (required for peptidomimetic pCoMole)
Peptidomimetic property oracles use the latest PeptiVerse SMILES predictors. Clone it next to this repo, or set PEPTIVERSE_ROOT:
# recommended: sibling of pCoMole
cd ..
git clone https://huggingface.co/ChatterjeeLab/PeptiVerse
cd PeptiVerse
pip install -r requirements.txt
# follow PeptiVerse README to download model weights
pCoMole looks for PeptiVerse in this order:
$PEPTIVERSE_ROOT../PeptiVerserelative to this repository
PeptiVerse scores the peptide-like side of each candidate. Admetica + DeepDTAGen still score the small-molecule side, and the original peptide-likeness mix logic is unchanged.
MAFFT (required for GFP pCoMole)
GFP excitation, brightness, and emission oracles run through FPredX and need MAFFT on your PATH.
# conda
conda install -c bioconda mafft
# or point to a local binary
export MAFFT_PATH=/path/to/mafft
Confirm with mafft --version before running GFP editing.
2. Train, evaluate, and sample Edit Flows
Run every command from the repository root.
SELFIES peptidomimetic Edit Flow
Training data is shipped at data/selfies/28k_mimetics.
# train
python train.py --config configs/config_selfies.yaml
# optional: python train.py --config configs/config_selfies.yaml --wandb
# evaluate a checkpoint on the validation split
python evaluate.py \
--config configs/config_selfies.yaml \
--ckpt peptidomimetics/ckpt/SELFIES_EditFlows.ckpt
# unconditional / seeded generation
python peptidomimetics/generate.py \
--config configs/config_selfies.yaml \
--ckpt peptidomimetics/ckpt/SELFIES_EditFlows.ckpt \
--input 'CSCC[C@H](NC(=O)[C@@H]1CCCN1C(=O)[C@H](Cc1ccc(O)cc1)NC(=O)[C@H](CCCNC(=N)N)NC(=O)[C@H](CO)NC(=O)[C@H](Cc1ccc(O)cc1)NC(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H]1CCCN1C(=O)[C@@H](N)CO)C(=O)N[C@@H](CC(=O)O)C(=O)O' \
--num_steps 30 \
--num_samples 2 \
--output_csv outputs/peptidomimetic_unconditional.csv
Or: bash peptidomimetics/scripts/train.sh and bash peptidomimetics/scripts/generate.sh.
GFP protein Edit Flow
A trained GFP checkpoint is shipped at gfp/ckpt/last_2.ckpt. The original GFP training arrows are not included. To retrain, put a HuggingFace load_from_disk dataset under data/gfp/{train,validation} (or edit configs/config_gfp.yaml).
python evaluate.py --config configs/config_gfp.yaml --ckpt gfp/ckpt/last_2.ckpt
python gfp/generate.py \
--config configs/config_gfp.yaml \
--ckpt gfp/ckpt/last_2.ckpt \
--input 'MSSGALLFHGKIPYVVEMEGNVDGHTFSIRGKGYGDASVGKVDAQFICTTGDVPVPWSTLVTTLTYGAQCFAKYGPELKDFYKSCMPDGYVQERTITFEGDGNFKTRAEVTFENGSVYNRVKLNGQGFKKDGHVLGKNLEFNFTPHCLYIWGDQANHGLKSAFKICHEITGSKGDFIVADHTQMNTPIGGGPVHVPEYHHMSYHVKLSKDVTDHRDNMSLKETVRAVDCRKTYDFDAGSGDTS' \
--num_steps 10 \
--num_samples 8 \
--output_csv outputs/gfp_unconditional.csv
3. pCoMole multi-objective editing
GFP
Official editor: gfp/pcomole.py (length + excitation + brightness, with GFP-classifier and emission constraints).
bash gfp/scripts/pcomole.sh
Equivalent command:
python gfp/pcomole.py \
--root_dir gfp/FPredX \
--config configs/config_gfp.yaml \
--ckpt gfp/ckpt/last_2.ckpt \
--num_steps 10 \
--num_candidates 50 \
--num_rollouts 10 \
--objective_weights 3 1 1 \
--output_file outputs/gfp_length_excitation_brightness.csv \
--input 'MSSGALLFHGKIPYVVEMEGNVDGHTFSIRGKGYGDASVGKVDAQFICTTGDVPVPWSTLVTTLTYGAQCFAKYGPELKDFYKSCMPDGYVQERTITFEGDGNFKTRAEVTFENGSVYNRVKLNGQGFKKDGHVLGKNLEFNFTPHCLYIWGDQANHGLKSAFKICHEITGSKGDFIVADHTQMNTPIGGGPVHVPEYHHMSYHVKLSKDVTDHRDNMSLKETVRAVDCRKTYDFDAGSGDTS'
If the output filename contains length_excitation_brightness, length_brightness, length_excitation, or length, the matching objective subset is used. Any other name uses all three objectives.
Peptidomimetics
Official editor: peptidomimetics/pcomole.py.
Objectives (7 scores when --specificity is on):
- non-toxicity
- solubility
- permeability
- half-life
- affinity
- motif
- specificity
PeptiVerse provides the peptide-side scores. Admetica + DeepDTAGen provide the small-molecule side. The two are mixed by peptide-likeness, same as the paper code. Motif / specificity still use the shipped BindEvaluator checkpoint.
# after PeptiVerse is cloned and its weights are downloaded
bash peptidomimetics/scripts/pcomole.sh
--target and --motifs are required for affinity and motif oracles. --specificity adds the seventh objective.
4. Included checkpoints
| Path | Role |
|---|---|
gfp/ckpt/last_2.ckpt |
GFP Edit Flow |
gfp/classifier_ckpt/best.pt |
GFP hard constraint |
gfp/FPredX/{ex,bright,em}_model |
FPredX property models |
peptidomimetics/ckpt/SELFIES_EditFlows.ckpt |
peptidomimetic Edit Flow |
peptidomimetics/ckpt/SMILES_BindEvaluator.ckpt |
motif / specificity |
peptidomimetics/ckpt/admetica/*.ckpt |
Admetica blending oracles |
peptidomimetics/ckpt/deepdtagen/ |
DeepDTAGen affinity |
Together these are about 10 GB. HuggingFace uploads should use Git LFS for the .ckpt / .pt / .pth files.
5. Environment variables
| Variable | Purpose |
|---|---|
PEPTIVERSE_ROOT |
PeptiVerse repo root (manifest + training_classifiers/) |
MAFFT_PATH |
MAFFT binary, if it is not on PATH |
GFP_CLASSIFIER_CKPT |
optional override for the GFP classifier |
6. Notes
- Run scripts from the repo root, or use the wrappers under
*/scripts/. - GFP FPredX will fail immediately if
mafftis missing; that is expected. - Peptidomimetic pCoMole will fail at oracle init if PeptiVerse is missing or its weights were not downloaded.
- Cas9 is not part of this release.
- Training writes to
outputs/by default.
Citation
If you use this code, please cite the pCoMole paper and PeptiVerse.