Reference recognizer, chembl_test_linux
Failure breakdown
Attempted: 6000 | exact: 3000 (0.500)
| failure mode |
n |
share of all |
share of failures |
what it points at |
| unparseable output |
0 |
0.000 |
0.000 |
decoder emitted a string RDKit cannot read |
| stereochemistry only |
0 |
0.000 |
0.000 |
skeleton correct; wedge/hash reading wrong |
| right formula, wrong bonds |
0 |
0.000 |
0.000 |
atoms counted correctly, connectivity misread |
| other structural error |
3000 |
0.500 |
1.000 |
a genuinely different molecule |
Accuracy by molecule size
| heavy atoms |
n |
exact |
valid |
| <=10 |
18 |
0.500 |
1.000 |
| <=15 |
168 |
0.500 |
1.000 |
| <=20 |
648 |
0.500 |
1.000 |
| <=25 |
1482 |
0.500 |
1.000 |
| <=30 |
1482 |
0.500 |
1.000 |
| <=35 |
1152 |
0.500 |
1.000 |
| <=40 |
654 |
0.500 |
1.000 |
| >40 |
396 |
0.500 |
1.000 |
Accuracy by rendering style
| style |
n |
exact |
valid |
| clean |
2000 |
0.500 |
1.000 |
| degraded |
2000 |
0.500 |
1.000 |
| varied |
2000 |
0.500 |
1.000 |
Sequence length
- Reference SMILES length, all attempted: median 47 tokens
- Reference SMILES length, correct predictions: median 47 tokens
Longest failures
| reference |
predicted |
O=S(=O)(O)c1ccc(-c2ccc(S(=O)(=O)Nc3cc(S(=O)(=O)O)cc4cc(S(=O)(=O)O)cc(S |
Nc1c(NCCOc2ccc(F)cc2-c2ccnc3[nH]c(C4=CCN(c5c(N)c(=O)c5=O)CC4)cc23)c(=O |
O=S(=O)(O)c1ccc(-c2ccc(S(=O)(=O)Nc3cc(S(=O)(=O)O)cc4cc(S(=O)(=O)O)cc(S |
CCCCOc1ccc(CCC(=O)O)cc1OCCCC |
O=S(=O)(O)c1ccc(-c2ccc(S(=O)(=O)Nc3cc(S(=O)(=O)O)cc4cc(S(=O)(=O)O)cc(S |
CCCCOC(=O)NS(=O)(=O)c1sc(CC(C)C)cc1-c1ccc(CN(C)C(C)=O)cn1 |
CC(=O)OC[C@H]1O[C@@H](O[C@H]2[C@H](OC(C)=O)[C@@H](OC(C)=O)[C@H](S(N)(= |
CCOP(=O)(OCC)C(NC(=S)NC(=O)[C@]1(C)CCC[C@]2(C)c3ccc(C(C)C)cc3/C(=N\O)C |
CC(=O)OC[C@H]1O[C@@H](O[C@H]2[C@H](OC(C)=O)[C@@H](OC(C)=O)[C@H](S(N)(= |
CCOC(=O)C1(Cc2ccc(OC)cc2)CCN(S(C)(=O)=O)CC1 |