CPPpred-En, a new ensemble machine learning tool, achieves 97.27% accuracy in predicting cell-penetrating peptides by combining protein language models with conventional peptide features — outperforming all existing prediction tools.
97.27% accuracyCPPpred-En achieved 97.27% accuracy in identifying cell-penetrating peptides — outperforming all existing prediction tools — by combining protein language model features with traditional peptide characteristics in an ensemble framework.
What the researchers found
CPPpred-En achieved state-of-the-art performance on two benchmark datasets:
- CPP924 dataset: 97.27% accuracy, MCC = 0.964
- MLCPP 2.0 dataset: 96.10% accuracy, MCC = 0.707
- Outperformed all existing prediction tools on both datasets
The key innovation was combining:
- Multiple protein language model (PLM) features (learned representations from large protein databases)
- Conventional peptide features (physicochemical properties, amino acid composition)
- Ensemble learning across multiple machine learning classifiers
- High-performing feature-classifier combinations selected and integrated
The model showed strong generalization across different datasets, demonstrating robustness.
Why it matters
Cell-penetrating peptides are crucial for the next generation of drug delivery — enabling therapies like siRNA, CRISPR components, and large molecules to enter cells that they otherwise couldn't reach. A prediction tool with 97% accuracy could save years of laboratory screening, allowing researchers to computationally identify the most promising CPP candidates before synthesizing and testing them. This accelerates the entire peptide-based drug delivery pipeline.
How the study worked
The researchers evaluated multiple types of features: protein language model (PLM) embeddings from several pretrained models and conventional peptide features (amino acid composition, physicochemical properties, etc.). These were tested across various machine learning classifiers. High-performing feature-classifier combinations were selected and integrated through ensemble learning. The model was trained and validated on two established CPP benchmark datasets (CPP924 and MLCPP 2.0) and compared against existing state-of-the-art prediction tools.
What this study cannot tell us
The model was validated computationally on existing datasets — no new experimental validation with synthesized peptides was performed. The MCC on the MLCPP 2.0 dataset (0.707) is notably lower than on CPP924 (0.964), suggesting variable performance across datasets. The model predicts binary CPP/non-CPP classification and may not capture nuances like cell-type specificity, uptake efficiency, or toxicity. Protein language models require significant computational resources. Performance on novel peptide classes not represented in training data is unknown.
How to read the evidence
This is a computational methodology study validated on benchmark datasets. While it demonstrates state-of-the-art prediction performance, it is an in silico tool that requires experimental validation. No new peptides were synthesized or tested. It represents a significant advance in computational peptide design tools.
When this study was published
Published in 2025, this study incorporates the latest protein language model technology, making it one of the most current AI tools for cell-penetrating peptide prediction.
The bigger picture
AI is transforming peptide science — from bioactive peptide discovery to property prediction to drug design. CPPpred-En represents the latest advance, leveraging protein language models (the peptide equivalent of large language models like GPT) to understand the 'grammar' of peptide sequences that determine cell-penetrating ability. As more peptide-based therapeutics enter clinical trials, computational tools like CPPpred-En become essential infrastructure for the field.
Questions still open
- How well does CPPpred-En perform on experimentally novel CPP sequences not represented in existing databases?
- Can the model be extended to predict not just whether a peptide penetrates cells, but how efficiently and through which mechanism?
- Would combining CPPpred-En predictions with structural modeling further improve the design of therapeutic cell-penetrating peptides?
Common questions
What are cell-penetrating peptides and why are they important?
How does AI help in peptide design?
Read the original research
CPPpred-En: Ensemble framework integrating a protein language model and conventional features for highly accurate cell-penetrating peptide prediction.
Computers in biology and medicine, 195, 110590
Citation
Jang, Yong Eun; Kwon, Minjun; Kim, Seok Gi; George, Nimisha Pradeep; Hwang, Ji Su; Basith, Shaherin; Lee, Gwang. (2025). CPPpred-En: Ensemble framework integrating a protein language model and conventional features for highly accurate cell-penetrating peptide prediction.. Computers in biology and medicine, 195, 110590. https://doi.org/10.1016/j.compbiomed.2025.110590