Reference recognizer, chembl_test_macfonts
Failure breakdown
Attempted: 6000 | exact: 3000 (0.500)
| failure mode |
n |
share of all |
share of failures |
what it points at |
| unparseable output |
0 |
0.000 |
0.000 |
decoder emitted a string RDKit cannot read |
| stereochemistry only |
0 |
0.000 |
0.000 |
skeleton correct; wedge/hash reading wrong |
| right formula, wrong bonds |
0 |
0.000 |
0.000 |
atoms counted correctly, connectivity misread |
| other structural error |
3000 |
0.500 |
1.000 |
a genuinely different molecule |
Accuracy by molecule size
| heavy atoms |
n |
exact |
valid |
| <=10 |
18 |
0.500 |
1.000 |
| <=15 |
168 |
0.500 |
1.000 |
| <=20 |
648 |
0.500 |
1.000 |
| <=25 |
1482 |
0.500 |
1.000 |
| <=30 |
1482 |
0.500 |
1.000 |
| <=35 |
1152 |
0.500 |
1.000 |
| <=40 |
654 |
0.500 |
1.000 |
| >40 |
396 |
0.500 |
1.000 |
Accuracy by rendering style
| style |
n |
exact |
valid |
| clean |
2000 |
0.500 |
1.000 |
| degraded |
2000 |
0.500 |
1.000 |
| varied |
2000 |
0.500 |
1.000 |
Sequence length
- Reference SMILES length, all attempted: median 47 tokens
- Reference SMILES length, correct predictions: median 47 tokens
Longest failures
| reference |
predicted |
O=S(=O)(O)c1ccc(-c2ccc(S(=O)(=O)Nc3cc(S(=O)(=O)O)cc4cc(S(=O)(=O)O)cc(S |
CC[C@@H](c1ccc(O)cc1)[C@H](CCCCn1cc(C(=O)O)c(=O)[nH]1)c1ccc(O)cc1 |
O=S(=O)(O)c1ccc(-c2ccc(S(=O)(=O)Nc3cc(S(=O)(=O)O)cc4cc(S(=O)(=O)O)cc(S |
OC[C@@H]1CCN(c2cc(-c3ccccc3)nc3ccccc23)C[C@@H]1O |
O=S(=O)(O)c1ccc(-c2ccc(S(=O)(=O)Nc3cc(S(=O)(=O)O)cc4cc(S(=O)(=O)O)cc(S |
O=C(C[C@@H](CCCC1CCCCC1)c1nc(CNS(=O)(=O)c2cccnc2)no1)NO |
CC(=O)OC[C@H]1O[C@@H](O[C@H]2[C@H](OC(C)=O)[C@@H](OC(C)=O)[C@H](S(N)(= |
CCOC(=O)c1ccccc1N[C@@H]1c2cc3c(cc2[C@@H](C2=CC(=O)C(=O)C(OC)=C2)[C@H]2 |
CC(=O)OC[C@H]1O[C@@H](O[C@H]2[C@H](OC(C)=O)[C@@H](OC(C)=O)[C@H](S(N)(= |
O=S(=O)(O)c1cc(Cl)c(O)c2ncc(Cl)cc12 |