MolScribe, novel scaffolds
Failure breakdown
Attempted: 900 | exact: 580 (0.644)
| failure mode |
n |
share of all |
share of failures |
what it points at |
| unparseable output |
81 |
0.090 |
0.253 |
decoder emitted a string RDKit cannot read |
| stereochemistry only |
101 |
0.112 |
0.316 |
skeleton correct; wedge/hash reading wrong |
| right formula, wrong bonds |
3 |
0.003 |
0.009 |
atoms counted correctly, connectivity misread |
| other structural error |
135 |
0.150 |
0.422 |
a genuinely different molecule |
Accuracy by molecule size
| heavy atoms |
n |
exact |
valid |
| <=15 |
6 |
0.833 |
1.000 |
| <=20 |
66 |
0.803 |
0.924 |
| <=25 |
243 |
0.728 |
0.942 |
| <=30 |
237 |
0.654 |
0.928 |
| <=35 |
174 |
0.569 |
0.874 |
| <=40 |
114 |
0.544 |
0.842 |
| >40 |
60 |
0.483 |
0.917 |
Accuracy by rendering style
| style |
n |
exact |
valid |
| clean |
300 |
0.827 |
0.987 |
| degraded |
300 |
0.347 |
0.797 |
| varied |
300 |
0.760 |
0.947 |
Sequence length
- Reference SMILES length, all attempted: median 47 tokens
- Reference SMILES length, correct predictions: median 45 tokens
Longest failures
| reference |
predicted |
Cc1c(C)c2c(c(C)c1OC(=O)CCC(=O)NC1CC(C)(C)N([O])C1(C)C)CCC(C)(C(=O)NC1C |
Cc1c(C)c2c(c(C)c1OC(=O)CCC(=O)NC1CC(C)(C)N(O)C1(C)C)CCC(C)(C(=O)NC1CC( |
C=C(C)[C@H]1C[C@@H](C[C@@]23C[C@@H](/C=C/C(C)C)C(C)(C)[C@@](CC=C(C)C)( |
C=C(C)[C@H]1C[C@@H](C[C@@]23C[C@@H](C=CC(C)C)C(C)(C)[C@@](CC=C(C)C)(C2 |
Cc1c(C)c2c(c(C)c1OC(=O)CCC(=O)NC1CC(C)(C)N([O])C1(C)C)CCC(C)(C(=O)NC1C |
Cc1c(C)c2c(c(C)c1OC(=O)CCC(=O)NC1CC(C)(C)N(O)C1(C)C)CCC(C)([C@@H](O)NC |
C=C(C)[C@H]1C[C@@H](C[C@@]23C[C@@H](/C=C/C(C)C)C(C)(C)[C@@](CC=C(C)C)( |
C=C(C)[C@H]1C[C@@H](C[C@@]23C[C@@H](C=CC(C)C)C(C)(C)[C@]2(CC=C(C)C)c2o |
Cc1c(C)c2c(c(C)c1OC(=O)CCC(=O)NC1CC(C)(C)N([O])C1(C)C)CCC(C)(C(=O)NC1C |
CC1=C(C)C(OC(C)C=CC(C)N[C@H]2CC(C)(C)[C@H](C)C2(C)C)[C@H](C)C=C1CC(C)( |