OMEGA is a computational discovery system. We don't ask for trust up front — this page shows exactly what we've already built from your group's own published work, with the same honesty standard we hold every model in our system to: no fabricated data, no unverified claims, held-out validation before anything is trusted.
The example below is real, unilateral work we've done on Prof. Ivan Aprahamian's published research (hydrazone photoswitch chemistry, Dartmouth) — this is a proposed, potential partnership, not an existing one. Nothing here is a projection — every number is from real papers or real calculations, independently re-verified before we trusted it.
We've built a structure-property regression model (Morgan fingerprint + scaffold-split validated — the same held-out methodology we use for every predictive model in our system, never a random split) and extended our DFT pipeline to handle isomerization-barrier scans specific to hydrazone switching. That model still correctly refuses to report an accuracy number against your own literature — the ask below hasn't changed: most of what we've mined has a measured property but no confirmed molecular structure yet, since every hydrazone scaffold in these papers lives in a Scheme or Figure, not in text a literature-mining pass can read.
The literature side kept growing after our first pass: the original 12 candidate papers, plus one more (the JACS 2025 steganography paper) read in full separately, plus two more real mining rounds (August 12 and 21) that found 6 additional papers — 19 total now — including your group's hydrazone/isosorbide chiral-dopant work, the synthetic negative-feedback-loop chemistry (with real SMILES and Zn(II)-affinity data extracted directly from its Supporting Information), and your most recent Science 2024 paper on a hydrazone-based molecular anion pump. The 6 new ones are real, genuine finds — but single-pass extractions, not yet held to the same independently-re-fetched standard as our first batch, and we're saying so plainly rather than blurring the two together.
On the computation side: we pointed our DFT pipeline at generating and validating candidate hydrazone structures ourselves — substituent variations around your published scaffolds — rather than only mining yours. As of this update, that's produced 25 real, DFT-relaxed candidates, each with its own confirmed structure and a computed switching barrier, thermal half-life, and absorption wavelength (22 predicted viable, 3 correctly flagged as non-viable), with more generated continuously. Separately, real GPU DFT isomerization-barrier scans are running against your own published hydrazone scaffolds — honestly, some of those runs have hit real technical walls (SCF convergence failures at specific geometries along the scan), which we're still working through, not papering over. None of this is a substitute for real measurements on your own compounds, and it's not yet enough data on its own for a defensible predictive model.
An update on the other two pillars: your group's own published research spans three real areas — hydrazone switches/motors (drug delivery, energy storage), hydrazone fluorophores (bio-imaging, sensing, OLEDs), and hydrazone-based sensing systems. Most of what's above is still the switching-kinetics side of the first. But we do now have a real, working computational screen for the fluorophore side too — seeded directly from your 2025 JACS super-resolution imaging paper, using real TD-DFT absorption calculations checked against the paper's own four known compounds before we trusted it. It honestly stops short of predicting real brightness (that needs excited-state fluorescence quantum yield, which this pass doesn't compute) and it's only screened a handful of real candidates so far — not a finished model, but a real, running one, not the empty gap we described before.
The sensing pillar has moved past "first screen built" to real results: we've now run real DFT-computed candidates through it, testing whether each one binds different real halides (fluoride, chloride, bromide, iodide) with meaningfully different strength at the same site — the first, necessary property any real chemical sensor needs. All of them clear that bar by a wide margin (real binding-energy spreads of 21-23 kcal/mol between the strongest and weakest halide), and the pattern is chemically consistent, not noise: fluoride binds far more strongly than the other three in every case, matching real halide chemistry (it's the smallest, most charge-dense halide, so it should be the strongest hydrogen-bond acceptor — and it is). That's still deliberately as far as it goes for now: knowing a receptor discriminates between halides is not the same as having a working sensor, which also needs a measurable optical or electrochemical signal tied to that binding, which we haven't built yet and are saying so plainly.
That's the one thing standing between "we've compiled your literature" and "we have a real, validated model that can rank compounds your group hasn't made yet." It's a small, specific ask — not a request for your whole lab notebook.
Specific, small, and scoped to what's genuinely missing — not an open-ended data request.
Your group's own compiled SI structures for the ~25 compounds we've already mined properties for (starting with the 10-compound JACS 2019 series) — a spreadsheet, a ChemDraw export, or the SI documents themselves. That alone unlocks a real, honest accuracy number for the model.
Optional, and only if it makes sense for your group: any measured switching data (thermal half-life, isomerization barrier, absorption/emission) that hasn't made it into a paper yet. This is the one thing a literature search can never reach, no matter how thorough — and it's exactly what a real collaboration can offer that publication-mining can't.
Our own internal roadmap for this work has 4 real stages: Level 1 is the structure-property corpus above; Level 2 is a machine-learned potential trained on real DFT data, letting fast calculations approximate the physics generally instead of one property at a time; Level 3 is a closed loop that generates new candidates and verifies them against that potential before anything is proposed. All three are computational, and all three stop at "here's what looks promising" — they cannot confirm a real compound actually behaves as predicted.
Level 4 is real synthesis and measurement — the field increasingly calls this a "self-driving laboratory" loop. We're honest that this is the one stage we cannot do ourselves: we have no robotic synthesis capability, and no substitute for a real chemist actually making a compound and measuring how it switches. If a real predictive model comes out of Steps 1–2 above and holds up on held-out data from your own literature, the natural next question is whether a handful of our computationally-proposed candidates are worth making — your group synthesizing them and measuring the real thermal half-life, isomerization barrier, or absorption/emission, and us feeding that real result back to recalibrate the model. A miss would be exactly as valuable as a hit here: either result is real information about whether computational screening works for this compound class at all.
To be direct about where this stands: this is the goal a potential partnership would be working toward, not something we're asking for today. Nothing past Step 1/2 above is a real ask until the model built from that data actually earns it.
Everything above is one specific, scoped piece of work on your group's hydrazone chemistry. It sits inside a larger system we've built for computational drug discovery and disease treatment more broadly — described below to the same honesty standard as everything above: real data, real held-out-validated predictors, and a plain statement of what hasn't been confirmed yet.
Same rule as above: a computational ranking is not a discovered drug. What we can say honestly is that real candidates have been generated against real data and are sitting in an outbound queue for each connected lab to look at — 24 for the GDSC line, 18 for CTRP, 12 for NCI-60, 18 for PRISM — not yet reviewed or confirmed by any of those groups.
Every one of these 12 connected data sources — including the work on your own lab above — now runs through two distinct, real search passes, not one: a narrow pass over the data source's own published/submitted results, and a separate, broader pass across the wider field (other groups' related work, competing methods, real-world clinical follow-up where it exists), kept in its own file with honest, never-auto-merged provenance. That broader pass isn't a one-time check either — it's re-verified on a standing 6-hour cadence, so a future data source added with no broad-search pass ever run becomes a visible, flagged gap instead of a silent one.
That's exactly where this stands today: real infrastructure, real data in, real validated predictors, real candidates out — and zero wet-lab or clinical confirmation yet anywhere in this broader system. We're disclosing that limit directly rather than letting the scale of the numbers above imply more than they've earned.
A real, held-out-validated structure-property model for your compound family, run against candidates your group hasn't synthesized yet — ranked, not fabricated, with an honest accuracy number attached. If the held-out accuracy is poor, that's the correct, useful outcome too: better to learn that from existing data than from a wasted synthesis effort. Nothing gets published or acted on without your review first.
Reach out however's easiest — email, a call, or just reply to whichever conversation brought you here.
[email protected]