The POSEIDON database and machine learning predictor can forecast cell-penetrating peptide uptake with 76% accuracy, accelerating drug delivery research.
r² = 0.76The machine learning model explains 76% of the variation in peptide cell uptake, a strong prediction for biological systems
What the researchers found
The POSEIDON database contains over 2,300 experimental entries with quantitative uptake values (how much peptide actually entered cells) and physicochemical properties for 1,315 peptides. Using this data plus genomic features of different cell lines, the team trained a regression model.
The model achieved a Pearson correlation of 0.87, Spearman correlation of 0.88, and r-squared of 0.76 on an independent test set. This means the model explains about 76% of the variation in peptide uptake across different cell types.
Why it matters
Getting drugs inside cells is one of the biggest challenges in medicine. Cell-penetrating peptides are a promising solution, but designing the right peptide for the right cell type has been largely trial and error. A tool that predicts uptake before any lab work could dramatically speed up drug delivery research.
The numbers in context
- 2,300+ experimental entries in the database
- 1,315 peptides with physicochemical profiles
- 1,200+ entries used for model training
- Pearson correlation: 0.87
- Spearman correlation: 0.88
- r² score: 0.76 on independent test set
How the study worked
Researchers curated a database from published experiments, standardizing uptake measurements across different studies. They then used over 1,200 entries with both peptide features and cell line genomic data to train a machine learning regression model. They validated it on an independent test set that the model had never seen during training.
Who was studied
Computational study analyzing 2,300+ experimental entries for 1,315 cell-penetrating peptides across multiple cell types
What this study cannot tell us
The database, while the largest of its kind, contains 2,300 entries from published literature, which may not cover all peptide types or cell types equally. Lab-to-lab variability in uptake measurements could introduce noise. The model was validated on held-out data but not prospectively tested in new experiments. Predicted uptake in cell culture may not translate to uptake in living organisms.
How to read the evidence
Rated moderate: well-validated computational tool tested on independent data, but predictions need experimental confirmation for new peptides.
When this study was published
Published in 2024. The database is freely accessible online and represents the state of the art in computational CPP prediction.
The bigger picture
Getting drugs inside cells is one of medicine's biggest challenges. A computational tool that predicts which peptide designs will be most effective at cell entry could dramatically speed up drug delivery research.
Questions still open
- Can the model predict uptake in cell types not in the training data?
- Will the predictions hold up in living organisms rather than cell culture?
Common questions
What is a cell-penetrating peptide?
How does the POSEIDON predictor work?
Read the original research
POSEIDON: Peptidic Objects SEquence-based Interaction with cellular DOmaiNs: a new database and predictor.
Journal of cheminformatics, 16(1), 18
Citation
Preto, António J; Caniceiro, Ana B; Duarte, Francisco; Fernandes, Hugo; Ferreira, Lino; Mourão, Joana; Moreira, Irina S. (2024). POSEIDON: Peptidic Objects SEquence-based Interaction with cellular DOmaiNs: a new database and predictor.. Journal of cheminformatics, 16(1), 18. https://doi.org/10.1186/s13321-024-00810-7