RSS Amplifier

The Curious Wavefunction · Jun 15, 2026

A sanity check on AI: Running a bioisostere replacement exercise for a complex macrocycle

0
Sign in to vote or save

Ash Jogalekar · The Curious Wavefunction

A good sanity check on an AI-assisted pipeline is when it does something that doesn’t look crazy or too creative, and when the system understands that when working on a challenging problem, conservative solutions are often the best. I tested this over the weekend on a challenging macrocyclic drug. The whole multistep pipeline took about 26 hours to set up and run.

The question as always was whether AI could run a pretty typical but useful modeling/med chem campaign: start from a real crystal structure, identify pharmacophoric interaction handles, propose conservative bioisosteres, generate full analogs, dock them back into the protein complex, check conformer accessibility, estimate local interaction energetics with quantum chemistry, run short MD pose-retention checks, compare properties, look at public/patent proximity, and then produce a report a chemist could actually read. And do all this with just a few prompts, with all the tools and code abstracted away and autonomously run.

The test case was daraxonrasib/RMC-6236, a large RAS(ON) macrocyclic inhibitor bound in a KRAS/CypA tri-complex crystal structure. For the purposes of this post, the shape of the problem is what was important: a big, chiral, heteroatom-rich macrocycle sitting in a ternary protein interface, with several obvious medicinal chemistry handles and a very crowded patent neighborhood.

Here is the parent structure used for the campaign.

The prompt was intentionally broad. I asked Codex to analyze the protein-ligand interactions in the crystal structure, estimate the major pharmacophoric interactions using a quantum chemical method such as SAPT, identify bioisosteric replacements, generate analogs, dock them, compare interaction energetics, check predicted properties, compare against patent/public-structure space, and build an HTML report plus notebook. I asked the system to carefully document its reasoning and limitations.

The final HTML report summarized the campaign this way:

Expanded analog generation, in-place Vina baseline, ETKDG conformers, SAPT0/jun-cc-pVDZ fragment energetics, properties, and patent proximity.

It also made the most important caveat explicit:

This is a computational prioritization report, not a legal freedom-to-operate opinion and not an experimental binding claim. SAPT values are fragment-pair estimates in the crystallographic/local pose context.

The parent molecule has several local features that are natural bioisostere targets. The campaign kept the overall scaffold constant and changed these features one or two at a time:

  • a thiazole-like heteroaryl handle

  • a chiral methoxy substituent

  • a piperazine tail

  • a pyridine/azine region

  • a macrocycle ester linker

  • a small cyclopropyl substituent

The replacements were deliberately conservative. This is important. The workflow was not a generative chemistry free-for-all. It was a pharmacophore-preserving analog campaign.

The core replacement logic looked like this:

Those are the kinds of changes a medicinal chemist would actually want to triage:

  • sulfur-to-oxygen in a five-membered heteroaryl ring

  • sulfur-to-NH heteroaryl polarity tuning

  • methoxy-to-fluoro as a small vector-preserving property change

  • methoxy-to-hydroxy as a polarity probe

  • N-methyl piperazine-to-piperidine as a tail lipophilicity/basicity probe

  • ester-to-amide as a macrocycle linker/stability probe

Codex generated 106 analogs from these single and paired replacements; it used existing literature references and known examples from databases like ChEMBL. It then ran a multi-stage triage:

  1. RDKit physicochemical property prediction

  2. AutoDock Vina score-only and local-only scoring in the parent binding site

  3. Local pose-retention RMSD

  4. ETKDGv3 conformer ensembles for prioritized analogs

  5. Psi4 SAPT0/jun-cc-pVDZ fragment interaction calculations

  6. Short explicit-solvent OpenMM MD pose-retention checks

  7. PubChem/ChEMBL exact-structure lookup

  8. Patent motif/proximity comparison

Once again, none of these tools had to be explicitly mentioned. The useful thing was the combination of these techniques. Docking alone is too fragile. Property filters alone are too generic. SAPT alone is too local. MD alone is too expensive and too easy to over-interpret. But together they form a reasonable computational triage funnel.

One thing I liked about the final report is that it made the work feel auditable. It was a sequence of ordinary computational chemistry tools glued together by the AI.

The timings below are reconstructed from output timestamps and MD logs, so they should be read as approximate wall-clock estimates rather than benchmark numbers. They were run locally on the available workstation environment (A MacBook Pro M5 with 32 GB RAM).

The punchline is that the “thinking” parts and the docking/property triage were fast. Unsurprisingly, the serious wall-clock cost came from the quantum chemistry and especially the MD. The 2 ns OpenMM runs reported about 8 ns/day on CPU, so they were perfectly reasonable as pose-retention sanity checks but not something to launch casually on every analog.

The best result was not exotic. It was the combination of two conservative changes:

thiazole to oxazole + methoxy to fluoro

That is a nice outcome. The model did not invent a bizarre new scaffold. It found that a smaller, less lipophilic, shape-preserving heteroaryl/fluoro pair looked attractive across several computational readouts. For such a complex macrocycle as this, you don’t want to propose large scaffold hops that will almost certainly change the binding mode in an uncertain manner.

Here is the parent compared with the best balanced analog. The changed pharmacophore atoms are highlighted.

The report card for this analog was:

Docking: local Vina -16.250 kcal/mol (-0.839 vs parent), pose RMSD 0.269 A. Conformers: best pose-like RMSD 0.868 A; minimum relative energy within 1.5 A = 0.000 kcal/mol. Properties: cLogP 5.465, TPSA 138.070, MW 782.962.

The parent local Vina score in the same setup was about -15.41 kcal/mol. The roughly 0.8 kcal/mol Vina difference in docking scores is statistically insignificant, but as a triage signal it was encouraging (again, the goal is to make conservative replacements). More importantly, the analog held the crystallographic pose during local minimization, had a low-energy ETKDG conformer close to that pose, reduced molecular weight, and did not damage the property profile. It’s the kind of safe analog a medicinal chemist would want to make, especially for a complex molecule like this one.

The leading analogs were variations on three medicinal chemistry themes: heteroaryl tuning, methoxy replacement, and tail/linker property probes.

The winning pattern was that the oxazole replacement preserved the heteroaryl vector while reducing sulfur-associated size/lipophilicity, and the fluoro replacement preserved a small substituent vector while trimming the methoxy group. In a big macrocycle, those are attractive changes because they are local enough not to destroy the binding mode.

The campaign used Psi4 SAPT0/jun-cc-pVDZ calculations on residue-centered, capped ligand fragments. I am not a quantum chemistry expert, but this seemed like a sensible compromise for a large ternary complex. Full quantum chemistry on the entire protein-ligand system is not realistic, and pure docking misses the physical character of local interactions.

The report described the SAPT setup like this:

SAPT was run with Psi4 SAPT0/jun-cc-pVDZ on residue-centered, capped ligand fragments. Odd-electron cuts were expanded to closed-shell fragments before retrying GLN63.

For the parent, the strongest reported local interaction was with a KRAS tyrosine contact, with total SAPT energy around -10 kcal/mol. Other important fragment contacts involved nearby KRAS and CypA residues, including methionine, glutamine, tryptophan, and phenylalanine contacts.

This is useful because it gives a sanity check on the analogs. If a proposed replacement improves a docking score but disrupts the interaction pattern that anchors the parent pharmacophore, it is less interesting. Conversely, a conservative replacement that preserves the interaction network and improves the property profile is worth keeping.

The oxazole/fluoro analog passed that smell test. It was not just a docking artifact; it preserved the local interaction logic well enough to remain the top candidate after the deeper ranking.

For large macrocycles, a docked pose can be misleading if the ligand cannot reasonably adopt that conformation. Codex therefore generated ETKDGv3 conformer ensembles for selected candidates and compared low-energy conformers with the crystallographic/local docked pose.

The best oxazole/fluoro analog had a pose-like conformer at essentially zero relative energy within the generated ensemble. That is one of the reasons it rose above some alternatives with good docking numbers but worse conformer penalties.

Four 2 ns explicit-solvent OpenMM simulations were run:

  • the parent

  • the oxazole/fluoro analog

  • the oxazole/hydroxy analog

  • the imidazole-NH/piperidine-tail analog

The report was careful about what this means:

These short trajectories are useful pose-retention checks, not converged binding free-energy simulations.

The parent and analogs stayed close to their starting poses. The parent mean ligand RMSD was about 1.10 Å. The selected analogs were in the same broad range, with mean ligand RMSDs of roughly 1.03 to 1.42 Å.

That does not prove binding. It does suggest that the top analog poses did not immediately fall apart in a solvated, restrained-protein simulation. For this kind of triage, that is a useful sanity check.

The best analog also looked reasonable by simple property metrics:

  • parent MW: about 811

  • best analog MW: about 783

  • parent cLogP: about 5.61

  • best analog cLogP: about 5.47

  • parent TPSA: about 134

  • best analog TPSA: about 138

So the top candidate did not “win” by simply becoming bigger and greasier. It got slightly smaller, kept lipophilicity in the same range, and only modestly increased polarity.

The patent/public-structure result was honest:

Exact InChIKey lookup found the parent in PubChem/ChEMBL. No top analog exact PubChem or ChEMBL record was found by InChIKey in this run. Motif-level proximity remains high because the downloaded WO families repeatedly cover related heterocycles and substituent classes.

That is exactly what one should expect. The top analogs were not exact public matches in this run, but they live close to a clinical-stage chemical series. They were flagged as high Markush/proximity risk. This campaign produced modeling hypotheses, not freedom-to-operate conclusions.

If I had to compress the campaign into one SAR hypothesis, it would be:

A heteroaryl sulfur-to-oxygen swap plus a methoxy-to-fluoro replacement is the most attractive conservative bioisostere pair for preserving the parent tri-complex pharmacophore while modestly improving the computational profile.

The secondary hypotheses are:

  • Imidazole NH can work geometrically but adds polarity/HBD risk.

  • Diazine tuning is plausible and worth keeping as an azine-vector probe.

  • Piperidine tail replacement can improve docking but risks cLogP creep.

  • Hydroxy substitution is geometrically tolerated but likely too polar.

  • Ester-to-amide is a reasonable stability/linker probe but not the top. computational hit.

The impressive part was not that an AI system knew some magic bioisostere; the replacements were standard medicinal chemistry moves. It proposed a few simple analogs and validated them across several computational checks. At the very least, these analogs now have enough confidence built into them that one could imagine running longer MD or FEP simulations.

But what was really impressive about the AI was the tool integration as well as the fact that it understood that for a system this complex, conservative substitutions are the best starting point:

  • It kept compound IDs synchronized across SDFs, CSVs, docking outputs, QM outputs, MD outputs, and reports.

  • It fixed chemistry and file-format issues as they appeared.

  • It generated full analog structures rather than abstract fragments.

  • It ran baseline parent scoring in the same protocol as the analogs.

  • It added conformer checks when the initial analysis was too shallow.

  • It did not overclaim patent novelty.

  • It produced a readable HTML report and notebook.

That is where a lot of real computational chemistry time goes and where these tools make a huge difference.

No posts

Read the original on medchemash.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.