• N&PD Moderators: Skorpio

How to Find Any Chemical Synthesis in Seconds (Free Database)

red22

Bluelighter
Joined
Nov 23, 2009
Messages
2,206
Video: h‍ttps://www.youtube.com/watch?v=lCJEr5bnjUU I embedded the video below.

Transcript: Hi everyone. If you're a chemist, you know the drill. You need to synthesize a specific molecule, but the biggest challenge isn't the lab work. It's finding a reliable protocol. Sure, you can try googling it, but that only works if you have a precise name. And even then, results are often hit or miss.

Today, I'm going to share my secret weapon with you, a massive searchable database of over 50,000 synthesis methods that I use every single day. It runs on Data Warrior, a powerful open-source program, and it's going to change how you do research. Here's how it looks. I use Data Warrior to open it. It's completely free. To find what you need, use the filter panel on the right.

Let's say we want to find Ditrophenol.
We draw the structure, click okay, and the list filters down.
Notice something. The search didn't just find the exact match. It performed a substructure search. This is a powerhouse feature that Google simply can't replicate. It allows you to find synthesis methods for similar molecules when an exact match isn't available.
However, if you need only denitrofenol, you can use the lasso tool, doubleclick the free atoms, and check prohibit further substitution.
Now only our target molecule remains.

Here is the literature reference. It's not clickable. From here you can open the folder with book. Here you can open specific pages.
You can open the links either from this table view by simply clicking on them or you can open them from right-click menu.
This isn't just a search engine. It's a massive digital library of verified lab manuals and gold standard procedures.
Just look at the list of literature.
Since this database was originally created for the postsviet scientific community, many books are in Russian.
But don't worry. First, about a third of the literature is already in English.
Like for example, this synthesis.

Second, for the Russian parts, you can simply take a screenshot of the page and feed it to Gemini or chat GPT. They will recognize the text and translate the protocol for you perfectly. It's not just for organic chemistry either. You can find inorganic methods too. Here is for example preparation of rainy nickel catalyst. Historically, this database was used in almost every Ukrainian research institute. It looked like this and ran on the now defunct ISIS base. In the original version, you had to manually find the files in folders. I've ported everything to Data Warrior and added clickable links to the PDFs. Just make sure to read the instructions in the description to make the links work.

You can download the whole database from my Google Drive. I will leave a link in the description. You might ask, "This database is 20 years old. Isn't there something newer online?" There are online services, but in my experience, none work as efficiently for finding actual procedures. First of all, you can try Google if you know the name of the substance. I should mention right away that it usually doesn't show the specific methods found in my database.
However, sometimes it might give you something new from recent articles or patents. Sometimes even better than what's in the database, sometimes not.

But it's always worth checking Google just in case. Then there are more specialized websites, but they all share one common problem. They either find nothing at all or they don't find a synthesis specifically, but rather every single article where the substance is mentioned. Finding an actual procedure among all those papers is nearly impossible. Take Chem Spider for example. It's a great tool if you need to find chemical properties. There is a structure search here, but let's go to the articles section. Look at how many there are. Good luck trying to find a synthesis method here. Organic synthesis is a good resource with only verified methods, but there aren't many of them.

It also has a structure search. Here is what it found for me. As you can see, there's no synthesis for what I'm looking for. This happens almost every time, which is why I don't really use it. Nevertheless, it's a noteworthy resource. I found some excellent protocols here in the past. Sure, Cambl is for searching molecules in patents.
Unfortunately, as we can see here, we have the same problem. It just throws in everything where the molecule is mentioned. Perhaps if the structure were more complex, the result would be more relevant. It's worth keeping in mind, but I personally don't use it.

If you need a deep dive for a PhD literature review, you'll need Reaxis or SciFinder, but those are incredibly expensive and usually only available to large universities. If you have access to Reaxis, you probably don't need this video. If you don't have Reaxis access, I'm working on another alternative, the Sprey database. It contains 14 databases and about 2.5 million reactions. It used to be available for download years ago, but now it's a paid service very similar to Reaxis. I'm currently merging these files and converting them into a modern format. When I finish, I'll post it on my Telegram channel. So, make sure to subscribe. If you find my work helpful, you can support me via buy me a coffee or my mono link. That's all for today.

Happy synthesizing and see you in the next


Description: database link in my telegram (google drive): h‍ttps://t.me/ninja_chemist/17
download everything from telegram + torrent file: h‍ttps://t.me/ninja_chemist/23


I downloaded the database and re-hosted it on MEGA: https://mega.nz/folder/h4BQhBAa#TrbO9FsdCgeUCuiP05_iWQ


 
Last edited:
confirmed working under linux

however, since you did this under windows your 7zip cell reference C:
jmc31337@laptop:~/Downloads/datawarrior610/datawarrior_linux/datawarrior$ ./datawarrior
ERROR StatusLogger Log4j2 could not find a logging implementation. Please add log4j-core to the classpath. Using SimpleLogger to log to the console...
Could not invoke browser: java.io.IOException: Failed to show URI:file:///C:/synthesis_data_base/SOP/Синтезы%20органических%20препаратов/sop_03/338-339.pdf

file:///C:/ ... for instance

however, if i manually peruse the pdf and hit gemini for translation it does indeed work as advertised
 
confirmed working under linux

however, since you did this under windows your 7zip cell reference C:
jmc31337@laptop:~/Downloads/datawarrior610/datawarrior_linux/datawarrior$ ./datawarrior
ERROR StatusLogger Log4j2 could not find a logging implementation. Please add log4j-core to the classpath. Using SimpleLogger to log to the console...
Could not invoke browser: java.io.IOException: Failed to show URI:file:///C:/synthesis_data_base/SOP/Синтезы%20органических%20препаратов/sop_03/338-339.pdf

file:///C:/ ... for instance

however, if i manually peruse the pdf and hit gemini for translation it does indeed work as advertised
By the way, "Синтезы органических препаратов" are just several collective volumes of Organic Syntheses translated into Russian. The original procedures are freely accessible at https://orgsyn.org .
 
Reaxys.

Not Reaxis.

So claiming similarity with a database that costs £40,000/annum to access but getting the name wrong I'm afraid does not inspire confidence.

Do I use Reaxys? Yes? Is this similar? No. Reaxys has access to 116 million documents. It also has far more sophisticated filters so finding the synthesis of specific moieties in the presence of other moieties takes about ten seconds. But if you use the ChemOffice suite, it automates so you draw a basic scaffold, it applies Markush structures (with filters), feeds the resulting SMILES string into Reaxys which has hyperlinks to the references.

For absolute free with no download needed and 119 million chemicals covered with the sole caveat that you need to provide the SMILES notation which even ChemSketch provides - PubChem.

Now something that would feed ChemOffice into PubChem would probably be of vast utility for people who are able to 'acquire' older versions of ChemOffice (12,13 or 14 are best) i.e. before they swapped to keeping the software on a remote-server and just provide high-speed connectivity. I assume this change was BECAUSE ChemOffice is one of the most pirated (in terms of value) suites of software on the planet.

Please don't think I'm asserting this has no value, only that it appears someone isn't being totally honest about the comparison and not for a moment do I think the OP is behind that.

FYI the best free workflow I know of is: ChemSketch to generate the SMILES string which you input to PubChem then use SciHub to locate the papers. It isn't AS efficient as the costly software BUT I would estimate that it provides around 70% of the utility and it's all well established software so we would know if any nasties were lurking.

In fact when a Russian paper deals with scaling they seem to apply a cost calculation based on the time, space, energy consumed, precursors, regents, solvents, co-reactants and give all that in an agreed unit so one can very quickly find the CHEAPEST route.

If someone could automate that calculation, that WOULD have real world value. In Russia it seems they still use pencil and paper to work it out. Which is fine - you cannot 'hack' a pencil and paper.
 
Last edited:
mispelling happens (i watched the video tutorial he posted he knows what he's doing)
However, my thing is why must i pay for these research papers, or why cant i just get the entire chemspider etc db they use (i have to go outta my way to format the query parameters or query strings OR use their website which is a drag )
They have their DB's on lockdown. Other way is to find reference manuals and index them

ie., i'm building Ai for chemistry searching

Code:
jmc31337@laptop:~/Downloads/chem-ai$ ollama serve

time=2026-10-03T14:32:39.845-04:00 level=INFO source=routes.go:2005 msg="server config" env="map[CUDA_VISIBLE_DEVICES: GGML_VK_VISIBLE_DEVICES: GPU_DEVICE_ORDINAL: HIP_VISIBLE_DEVICES: HSA_OVERRIDE_GFX_VERSION: HTTPS_PROXY: HTTP_PROXY: LLAMA_ARG_FIT: LLAMA_ARG_FIT_TARGET: NO_PROXY: OLLAMA_CONTEXT_LENGTH:0 OLLAMA_CREATE_REMOTE:false OLLAMA_DEBUG:INFO OLLAMA_DEBUG_LOG_REQUESTS:false OLLAMA_EDITOR: OLLAMA_FLASH_ATTENTION:false OLLAMA_GO_TEMPLATE:true OLLAMA_GPU_OVERHEAD:0 OLLAMA_HOST:http://127.0.0.1:11434 OLLAMA_IGPU_ENABLE: OLLAMA_KEEP_ALIVE:5m0s OLLAMA_KV_CACHE_TYPE: OLLAMA_LLM_LIBRARY: OLLAMA_LOAD_TIMEOUT:5m0s OLLAMA_MAX_LOADED_MODELS:0 OLLAMA_MAX_QUEUE:512 OLLAMA_MAX_TRANSFER_STREAMS:4 OLLAMA_MODELS:/home/jmc31337/.ollama/models OLLAMA_NOHISTORY:false OLLAMA_NOPRUNE:false OLLAMA_NO_CLOUD:false OLLAMA_NUM_PARALLEL:1 OLLAMA_ORIGINS:[http://localhost https://localhost http://localhost:* https://localhost:* http://127.0.0.1 https://127.0.0.1 http://127.0.0.1:* https://127.0.0.1:* http://0.0.0.0 https://0.0.0.0 http://0.0.0.0:* https://0.0.0.0:* app://* file://* tauri://* vscode-webview://* vscode-file://*] OLLAMA_REMOTES:[ollama.com] OLLAMA_SCHED_SPREAD:false OLLAMA_VULKAN:true ROCR_VISIBLE_DEVICES: http_proxy: https_proxy: no_proxy:]"

....

[GIN] 2026/10/03 - 14:46:27 | 200 |         2m41s |       127.0.0.1 | POST     "/api/embed"


Code:
jmc31337@laptop:~/Downloads/chem-ai$ python3 -m venv venv
source venv/bin/activate
(venv) jmc31337@laptop:~/Downloads/chem-ai$ find scripts
scripts
scripts/__pycache__
scripts/__pycache__/websearch.cpython-314.pyc
scripts/websearch.py
scripts/index.py
scripts/chat.py
(venv) jmc31337@laptop:~/Downloads/chem-ai$ find knowledge
knowledge
knowledge/datasets
knowledge/papers
knowledge/papers/The Psilocybin Producers Guide (Adam Gottlieb).pdf
knowledge/papers/psilocynextraction.pdf
knowledge/papers/Psilocybin_Patent_EtOH_Water_Extraction.pdf
knowledge/papers/pharmaceuticals-psilocybin.pdf
knowledge/books
knowledge/notes
knowledge/notes/ChemistryRXNS.txt
knowledge/notes/psilocybin_excerpt.txt
knowledge/notes/Preparing_an_Exceptional_THCA_tea.txt
(venv) jmc31337@laptop:~/Downloads/chem-ai$ python3 scripts/chat.py --source=www
Chemistry AI ready. Type 'exit' to quit.
Active source: www

Chemistry > whats the smiles code for caffeine?

Searching web...

Web sources:
FETCH ERROR: 403 Client Error: Forbidden for url: https://www.researchgate.net/figure/Chemical-structure-of-caffeine-C-8-H-10-N-4-O-2-CAS-58-08-2-SMILES_fig1_351130458
- Caffeine - Wikipedia
  https://en.wikipedia.org/wiki/Caffeine
Skipping blocked domain: https://pubchem.ncbi.nlm.nih.gov/compound/Caffeine
- InChI Key Database ⚛️ | Caffeine
  https://www.inchikey.info/ikdb/rz/t/rztamfziaatzdj-oahllokosa-n
- CAS, SMILES, InChl and Fingerprints - Personal page of Zhilong Jia
  https://zhilongjia.github.io/posts/2021/06/SMILES/
- Chemical structure of caffeine (C 8 H 10 N 4 O 2; CAS 58-08-2; SMILES ...
  https://www.researchgate.net/figure/Chemical-structure-of-caffeine-C-8-H-10-N-4-O-2-CAS-58-08-2-SMILES_fig1_351130458
  [no text extracted]
- Caffeine | C8H10N4O2 - ChemSpider
  https://www.chemspider.com/Chemical-Structure.2424.html
- Caffeine Structure, Formula, Molecular Weight, SMILES & 3D Viewer
  https://www.moldraw.com/blog/molecules/caffeine.html
- When are Two Compounds the Same? The effect of SMILES format on ...
  https://www.biosymetrics.com/blog/representing-similar-compounds

DEBUG RAW:
{
  "smiles": "CN1C=NC2=C1C(=O)N(C(=O)N2C)C"
}

Qwen:

CN1C=NC2=C1C(=O)N(C(=O)N2C)C

Evidence verification:
NOT AVAILABLE - model did not return structured evidence.

Sources:
[SOURCE 1] Caffeine - Wikipedia
https://en.wikipedia.org/wiki/Caffeine
[SOURCE 2] InChI Key Database ⚛️ | Caffeine
https://www.inchikey.info/ikdb/rz/t/rztamfziaatzdj-oahllokosa-n
[SOURCE 3] CAS, SMILES, InChl and Fingerprints - Personal page of Zhilong Jia
https://zhilongjia.github.io/posts/2021/06/SMILES/
[SOURCE 4] Caffeine | C8H10N4O2 - ChemSpider
https://www.chemspider.com/Chemical-Structure.2424.html
[SOURCE 5] Caffeine Structure, Formula, Molecular Weight, SMILES & 3D Viewer
https://www.moldraw.com/blog/molecules/caffeine.html
[SOURCE 6] When are Two Compounds the Same? The effect of SMILES format on ...
https://www.biosymetrics.com/blog/representing-similar-compounds

Chemistry >

or once i grab bunch books and knowledge database i can locally ONLY search a particular pdf

Code:
(venv) jmc31337@laptop:~/Downloads/chem-ai$ python3 scripts/index.py
/home/jmc31337/Downloads/chem-ai/scripts/index.py:4: DeprecationWarning: `langchain-community` is being sunset and is no longer actively maintained. See https://github.com/langchain-ai/langchain-community/issues/674 for details and migration guidance toward standalone integration packages.
  from langchain_community.document_loaders import (
Scanning: /home/jmc31337/Downloads/chem-ai/knowledge
PDF: /home/jmc31337/Downloads/chem-ai/knowledge/papers/The Psilocybin Producers Guide (Adam Gottlieb).pdf
PDF: /home/jmc31337/Downloads/chem-ai/knowledge/papers/psilocynextraction.pdf
PDF: /home/jmc31337/Downloads/chem-ai/knowledge/papers/Psilocybin_Patent_EtOH_Water_Extraction.pdf
PDF: /home/jmc31337/Downloads/chem-ai/knowledge/papers/pharmaceuticals-psilocybin.pdf
TEXT: /home/jmc31337/Downloads/chem-ai/knowledge/notes/ChemistryRXNS.txt
TEXT: /home/jmc31337/Downloads/chem-ai/knowledge/notes/psilocybin_excerpt.txt
TEXT: /home/jmc31337/Downloads/chem-ai/knowledge/notes/Preparing_an_Exceptional_THCA_tea.txt
Documents loaded: 72
Chunks created: 132
Removing old database...
Creating vector database...
Finished.
Stored in: /home/jmc31337/Downloads/chem-ai/embeddings

Indexed sources:
 - ChemistryRXNS.txt
 - Preparing_an_Exceptional_THCA_tea.txt
 - Psilocybin_Patent_EtOH_Water_Extraction.pdf
 - The Psilocybin Producers Guide (Adam Gottlieb).pdf
 - pharmaceuticals-psilocybin.pdf
 - psilocybin_excerpt.txt
 - psilocynextraction.pdf
(venv) jmc31337@laptop:~/Downloads/chem-ai$

and pass python3 --source=psilocybin_excerpt.txt searching ONLY that one db knowledge set


btw my Ai got the smiles right off paralell 10 DuckDuckGo parsing sites (avoiding .gov .mil endpoints which fucks my access over to other pubchem gov ran sites)

so back to the point, Ai is gettin great and makes research a lot easier and personal Ai is uncensored
 
I know nothing about chemistry, but think it's so cool that you guys could potentially make any of these drugs i take, if given the right chemicals needed. As someone who takes antidepressants anxiety medication, and preventative meds like subs it's truly amazing to understand what chemicals make your brain do what. Kudos to all of you
 
That text was automatically generated by Tactiq.

Well, I feel that suggests use of that tool may not provide reliable results.

I hate to say it but you cannot place blame on a computer program which is why man-in-the-loop is so vital.

As I pointed out - the way Russian papers cost out a synthesis IS something of real value and if someone is intent on developing tools for the database, automation of this calculation would be a very good one.
 
How do you get into Reaxys, doesn't it need a username (email) and password? I doubt that it's just for free, it looks like a premium service.
 
It IS a premium service. It certainly isn't free.
How are you getting it then? Do you have to pay for it personally out your own pocket? I'm assuming you have contacts that gave you access free of charge? Assuming one was going to pay for it out their own pocket have you any idea how much it would cost? I'm assuming it's thousands.
 
Well, obviously detailing the loop-hole would see that loop-hole closed.

I think what I can tell you is that if a large institution wants a licence for a large number of users, there is no formula. It's negotiated. So an extra user can be added semi-officially AFTER the licence has been agreed upon simply because RELX sees the cost of re-negotiation as more time and effort and crucially - means they don't get the cash as quickly.

No magic involved. Simply staying in touch with one's alma mater.

I didn't ask, it was offered.

I just end up having to sort of be a 'very helpful indeed' peer reviewer from time to time. Not an advisor and not a peer reviewer, but more like a crammer.

In short, old boys network.
 
Top