Documentation
How to use the Neuralocity catalog, and what the drug-discovery concepts behind every field and score on this platform actually mean.
Workspace Overview
The Neuralocity catalog holds several million computationally generated candidate molecules. Every molecule has been scored against a set of biological target classes and a set of safety/interaction models, so you can screen for promising candidates before any wet-lab work begins.
The top navigation bar gives you access to the Catalog (browse and filter molecules), your Cart (molecules you’ve selected), this Documentation page, and your account settings (gear icon, top right).
The Catalog Page
The Catalog page is a filterable, sortable table of every molecule in the current release. Each row shows a compact set of columns; the full detail for a molecule lives on its own page (see below).
Filters
- Target Class + Mechanism: choose a biological target class and whether you want antagonist or agonist activity. This is required before the Activity column has any value to show or sort by.
- Activity Tier: restrict results to one or more tiers (A through E) once a target class is selected.
- Safety: restrict to molecules classified Safe or Not Safe.
- Drug-Drug Interaction (DDI) flags: restrict by how many CYP450 isoforms a molecule is predicted to inhibit.
- Property ranges: Molecular Weight, LogP, H-Bond Donors/Acceptors, QED, and SA Score all support min/max range filters.
- Text search: search by identifiers such as Molecule ID or SMILES.
- Similarity search: paste a SMILES string and a similarity threshold to find structurally similar molecules already in the catalog (see Molecular Similarity Search).
Reading the table
Columns are driven by what’s actually available for a molecule: a column showing “Not available” usually means that data hasn’t been populated for that row yet, or (for Activity) that no target class filter is selected. Click the column header to sort; click again to reverse direction.
Every row has a Details link (and the Molecule ID itself is also a link) that opens the full molecule detail page, and a cart button to add that molecule directly. You can also select all molecules matching your current filters and add them to your cart in bulk.
The Molecule Detail Page
Clicking into a molecule shows everything known about it, organized into panels:
- Structure: a rendered image of the molecule. Click it to open a larger version.
- Identity: the molecule’s identifiers: SMILES, Canonical SMILES, SMARTS, InChI, InChI Key, and Formula (see Molecular Identity).
- Properties: stored physicochemical properties plus structural characteristics computed on the fly (see Physicochemical Properties and Structural Characteristics).
- Activity Tiers, Safety, and Drug-Drug Interaction: the same scoring concepts as the Catalog page’s filters, shown in full for this molecule.
- Synthesizability & Stability: an AI-based assessment of how practical the molecule is to make and how well it would hold up in storage (see Synthesizability & Stability).
- Predict Binding Affinity For Your Target: paste your own target sequence to get a computational predicted structure for this molecule against it (see Predicted Structure & Binding Affinity).
- Provenance: the molecule’s source identifier within the catalog.
Cart & Purchases
Add molecules to your cart from the Catalog page or a molecule’s detail page. Adding a molecule to your cart claims it, so it drops out of the catalog for everyone else until you remove it or purchase it. Your cart persists across sessions. From the Cart page, there are two ways to act on it:
- Buy digital listing: an instant, fixed-price purchase of that molecule’s structure and property data. Checkout is handled by Stripe, and once payment completes, purchased molecules move from your Cart to your Purchase History page under Account settings, where you can download the data as CSV.
- Synthesize & deliver: request to have the physical compound synthesized and shipped to you. Pricing varies per molecule, so this doesn’t go through checkout. Clicking “Request a quote” emails your cart contents to our team, and we follow up directly with a firm quote and an expected delivery window.
Account & Organization
Your account settings (gear icon, top right) let you manage your profile, view your organization’s members, send invites if your organization supports multiple users, and review your purchase history.
Some fields on the platform may be locked behind an entitlement (indicated with a lock icon), depending on your organization’s plan.
Drug Discovery Concepts in This Platform
Every concept below corresponds to a real field or score you can see somewhere in the product. Nothing here is generic background that isn’t actually represented in the catalog.
Molecular Identity
A single molecule can be written down in several different, standardized ways. The Identity panel on a molecule’s detail page shows all of them:
- SMILES (Simplified Molecular Input Line Entry System): a compact text notation for a molecule’s structure. It’s the format the molecule was originally generated in.
- Canonical SMILES: the same molecule, rewritten into one standardized SMILES form. Two chemically identical molecules can be written as different SMILES strings; canonicalization gives every molecule one consistent representation, which is what deduplication and similarity search rely on.
- SMARTS: a related notation used for describing structural patterns rather than a single specific molecule (e.g. “any aromatic ring with an attached carbonyl”). Useful for substructure searching.
- InChI (International Chemical Identifier) and InChI Key: a different standardized identifier format, widely used across chemistry databases. The InChI Key is a short, fixed-length hash of the full InChI, convenient for lookups and cross-referencing against other databases.
- Formula: the molecular formula (e.g. C₁₃H₁⁾N₂O₂), a simple count of each element present.
Physicochemical Properties
These are computed once, when a molecule enters the catalog, and stored.
- Molecular Weight (MW): the mass of the molecule. Most successful oral small-molecule drugs fall in a fairly narrow weight range, roughly 150 to 500 daltons, and MW is one of the first things medicinal chemists check because it touches nearly every other property that matters:
- Absorption and permeability. Larger molecules generally cross cell membranes and the gut wall less efficiently, making it harder for an oral drug to reach the bloodstream in useful amounts.
- Distribution. Once absorbed, a molecule’s size affects how easily it moves through tissue and reaches its target, and how quickly the body clears it.
- Binding specificity. Very small molecules often lack enough surface area to bind a target tightly or selectively, which can mean weaker potency or more off-target effects.
- Synthesis and formulation. Heavier, more complex molecules typically take more synthetic steps to build and can be harder to formulate into a stable, dosable product.
- Dosing practicality. A heavier molecule generally means more milligrams are needed per dose to reach an effective concentration in the body.
- LogP: a measure of lipophilicity, how a molecule partitions between a fat-like phase and water. This is the property on this platform most directly relevant to a molecule’s solubility and membrane-permeability trade-off: too low (too water-loving) and a molecule struggles to cross cell membranes, while too high (too fat-loving) it tends toward poor solubility and off-target binding.
- H-Bond Donors and H-Bond Acceptors: counts of atoms that can donate or accept a hydrogen bond. Along with Molecular Weight and LogP, these are the four properties behind Lipinski’s Rule of Five (see below).
- QED (Quantitative Estimate of Drug-likeness): a single 0-to-1 score that blends several drug-likeness properties into one number. Closer to 1 means the molecule’s overall property profile looks more like known, successful drugs.
- SA Score (Synthetic Accessibility Score): a computed 1-to-10 score based purely on structural complexity, how unusual or intricate the molecule’s substructures are compared to fragments that appear often in known, easily-made molecules. It does not reason about an actual route to make the molecule; it is a fast heuristic for how complicated the structure looks. This is different from the AI-based Synthesizability rating and description described later, which reason about a specific, plausible synthetic route (how many steps, what reagents, what could go wrong) to assess whether the molecule could actually be made. A structurally simple-looking molecule can still get a difficult Synthesizability rating if a realistic route is genuinely hard, and vice versa, so the two may disagree; both are shown because they capture different things.
- Predicted Solubility (Log S): predicted aqueous solubility, as log₁₀ of the molecule’s solubility in water in moles per liter. Unlike the properties above, this is not computed directly from the structure by a fixed formula; it is predicted by a trained graph neural network (AttentiveFP) that learned the relationship between structure and measured solubility. Higher (less negative) values mean more soluble. Below −4 is a widely used rule-of-thumb cutoff in early drug discovery for compounds likely to cause problems in aqueous assays; above that, more solubility is not itself a sign of a better candidate, just an increasing degree of the same acceptable property, which is why the catalog colors it as shades of blue above the cutoff rather than a best-to-worst gradient.
Structural Characteristics
Shown on a molecule’s detail page, computed directly from its structure each time the page loads.
- Heavy Atoms, Hydrogens, and Total Atoms: basic atom counts (heavy atoms are every atom except hydrogen; total atoms is heavy atoms plus hydrogens).
- Heteroatoms, O + N Atoms, Halogen Atoms, and Non-Organic Atoms: counts of atoms other than carbon, broken down by category. A high halogen count, for example, can be a flag for metabolic-stability or toxicity risk in medicinal chemistry.
- Fragments: the number of disconnected pieces in the structure. Most drug candidates should be a single connected fragment.
- Rings and Rotatable Bonds: ring count and the number of single bonds that can freely rotate. Rotatable bond count is a common proxy for molecular flexibility, which affects both binding and oral bioavailability.
- Stereocenters (defined / unspecified / unknown) and Cis/Trans Bonds (defined / unspecified / unknown): a stereocenter is an atom whose spatial arrangement can produce distinct molecules (mirror-image or otherwise) with potentially very different biological activity. Because these molecules come from a generative process rather than a specific experimental synthesis, many stereocenters are left unspecified in the original structure. The count tells you how much stereochemical ambiguity a given molecule has.
- Polar Surface Area (PSA): the surface area of the molecule contributed by polar atoms. A widely used predictor of a molecule’s ability to cross the blood-brain barrier and its general absorption behavior.
- Molar Refractivity: a measure related to a molecule’s size and polarizability, another property that correlates with drug-likeness.
- Rule of Five Violations: Lipinski’s Rule of Five is a widely-used rule of thumb: a molecule is more likely to be orally bioavailable if it has Molecular Weight ≤ 500, LogP ≤ 5, H-Bond Donors ≤ 5, and H-Bond Acceptors ≤ 10. This count is simply how many of those four thresholds the molecule fails, where 0 means it satisfies all four.
Target Activity & Tiers
Every molecule is scored against ten biological target classes, grouped by biological classification:
- Kinase
- GPCR
- Protease
- Ion Channel
- Nuclear Receptor
- Enzyme
- Transporter
- Epigenetic
- Apoptosis Regulator
- Aggregation Target
Each target class is scored separately for antagonist activity (blocking/inhibiting the target) and agonist activity (activating/stimulating the target), since a molecule’s predicted behavior can differ substantially between the two mechanisms. Each target class/mechanism pair is its own machine learning model, trained on known compound activity data for that target class. The output is a model prediction, not a measured result.
Activity is shown as a letter tier from A (best) to E (worst):
- A: highest-confidence predicted activity.
- B, C, D: progressively lower predicted activity.
- E: lowest predicted activity, or not enough signal to place the molecule in a higher tier.
Model Methodology
Each target class/mechanism activity model is trained on known compound activity data (drawn from ChEMBL, a public repository of experimentally measured compound-target interactions) and evaluated on a held-out test set that is deliberately hard.
Scaffold-clustered holdout
Before training, every compound is grouped into a cluster by its core scaffold, the ring system at the heart of the molecule, independent of whatever side chains are attached. Two molecules that differ only in a side chain end up in the same cluster. Roughly 70% of these scaffold clusters are set aside for training and 30% are held out entirely for testing. Because the split happens at the cluster level rather than the individual-molecule level, no scaffold appears in both the training set and the test set. The model’s reported accuracy comes from scoring it against that 30%, so it reflects performance on chemically unfamiliar scaffolds, not just unseen molecules that happen to be close variants of something already seen during training.
Comparison to a random split
A simpler random split, holding out a random 30% of individual molecules, can let a model look better than it really is: near-identical analogs of a training molecule frequently end up in a random test set too, so the model can effectively memorize a scaffold rather than learn to recognize activity for chemistry it hasn’t encountered. That inflates reported performance in a way that doesn’t reflect how well the model generalizes to a genuinely new molecule. The scaffold-clustered holdout used here is a stricter, more realistic test.
As with every score, even a well-designed holdout evaluation is still a computational estimate of performance, not a substitute for experimental validation.
Predicted Solubility (Log S) model
Unlike the target-class and toxicity models above, which train a separate classifier per biological question on a fixed molecular fingerprint, solubility uses a single graph neural network (an architecture called AttentiveFP) that reads a molecule’s atoms and bonds directly rather than reducing it to a fingerprint first, and predicts a continuous number (log S) rather than a class probability.
Training data comes from solubility datasets, built from experimentally measured aqueous solubility values (not computed or predicted). The model is first pre-trained on the larger collection for broader chemical coverage, then fine-tuned, the same benchmark used to report accuracy: on a held-out scaffold split (the same scaffold-clustered methodology described above, so unseen chemistry, not just unseen molecules) it reaches a root-mean-square error of about half a log unit and a Spearman correlation of 0.96 with the true measured values.
This model predicts a general physicochemical property, not activity against a specific biological target, so it is shown as its predicted value directly rather than converted into a tier. As with every score on this platform, it is a computational estimate, not a lab measurement.
Safety & Toxicity
Every molecule is screened against eight toxicity endpoints, each its own trained machine learning model:
- Hepatotoxicity: liver toxicity risk.
- Carcinogenicity: cancer-causing potential.
- Cardiotoxicity: heart toxicity risk.
- Gastrotoxicity: gastrointestinal toxicity risk.
- Neurotoxicity: nervous-system toxicity risk.
- Psychiatric toxicity: risk of adverse psychiatric effects.
- Respiratory toxicity: respiratory-system toxicity risk.
- Teratogenicity: risk of causing developmental/birth defects.
These eight predictions are combined into one weighted Safety indicator, shown as a simple Safe / Not Safe classification on the Catalog page (a green or red dot) and in full detail on the molecule page. Endpoints with the highest real-world consequence are weighted most heavily in that combined score. As with Activity, this is a computational prediction intended for triage, not a substitute for actual toxicology testing.
Drug-Drug Interaction (DDI) Screening
Many drugs are broken down in the body by a family of liver enzymes called CYP450 (cytochrome P450). If a new molecule strongly inhibits one of these enzymes, it can interfere with how other, unrelated drugs that rely on that same enzyme get metabolized. This is a common real-world cause of drug-drug interactions.
Every molecule is screened for predicted inhibition of five CYP450 isoforms:
- CYP1A2
- CYP2C9
- CYP2C19
- CYP2D6
- CYP3A4/5
The DDI indicator shown on a molecule is a count from 0 to 5: how many of these five isoforms the molecule is predicted to meaningfully inhibit. This is a proxy for interaction risk through shared metabolic pathways, not a lookup of specific known drug-vs-drug interactions.
On the molecule detail page, each flagged isoform also lists the known drugs that rely on it for clearance: the specific medications that could build up to unsafe levels if co-administered with this molecule.
Predicted Structure & Binding Affinity
This panel, on the molecule detail page, is different from every other score on this platform: instead of a pre-computed result, you supply the target. Paste a protein, DNA, or RNA sequence for a target you’re interested in, and the platform runs a computational structure prediction for this molecule against that specific target, something the rest of the catalog’s pre-computed scores don’t do, since those are scored against target classes, not a target you provide yourself.
This can take some time to run. A progress indicator shows while it works, and you can navigate away and come back later; the result will be waiting for you.
Reading the result
- 3D structure: the target is shown as a cartoon ribbon, your molecule as thick orange sticks, and nearby target residues (the local binding pocket) as thin gray sticks for context. Drag to rotate, scroll to zoom, and use the button below the viewer to switch between a close-up of the binding site and the full structure.
- Confidence (pLDDT): the model’s overall confidence in the predicted structure as a whole, on a 0-to-1 scale.
- Interface Confidence: the model’s confidence specifically in the predicted contact between your molecule and the target, shown as High, Moderate, or Low. This is separate from the overall structure confidence above, and is the closest indication currently available of whether the predicted pose looks plausible.
- Predicted IC50 and Binding Affinity Likelihood: a direct measure of predicted binding strength. These only appear if you check “Also predict binding affinity” before submitting; this step is significantly slower than the structure prediction above, so it is opt-in rather than always run.
This is a computational prediction, not an experimentally validated result. The model always produces a predicted structure for whatever target and molecule you give it; it does not currently have a way to say “this molecule doesn’t bind.” A low Interface Confidence is the best available hint that a predicted pose may not be reliable, but it is not the same as a formal binding probability.
Refine with AutoDock Vina
Once a structure prediction completes, you can optionally run a second, independent check on it: AutoDock Vina, a classical physics-based docking method, searches for the best-fitting position and orientation of your molecule within the predicted pocket. This is a genuinely different approach from Boltz-2’s own deep-learning-based prediction above, not a re-run of the same method, so agreement or disagreement between the two is meaningful.
- Vina Binding Affinity: Vina’s own predicted binding strength, in kcal/mol. This is a different scale than Predicted IC50 above and the two are not directly comparable to each other.
- Pose Agreement (RMSD vs. Boltz): the distance, in Angstroms, between where Vina independently placed your molecule and where Boltz predicted it. A low value means the two methods agree on the binding pose; a high value means they disagree on where or how the molecule binds, worth noting either way.
- 3D overlay: Boltz’s predicted pose is shown in orange, and Vina’s independently-docked pose is shown in teal, on the same structure, so the two can be compared visually rather than only as a number.
This refines a predicted structure, not a validated one. Vina is scoring the same computationally-predicted pocket Boltz produced, not a crystal structure, so agreement between the two methods means they converge given that starting point, not that either is confirmed correct.
Attribution
This feature is powered by Boltz-2 (Passaro, Corso, Wohlwend, et al.), an open-source biomolecular structure and binding-affinity prediction model released under the MIT License, and uses the ColabFold server (Mirdita et al.) for sequence alignment. The refinement step above is powered by AutoDock Vina (Trott & Olson, 2010), molecular docking engine.
Synthesizability & Stability
Unlike every score described above, this panel (on the molecule detail page) is an AI-based opinion. It gives a written assessment and a 1-to-6 rating on two independent axes:
Synthesizability
How practical the molecule looks to actually make, reasoning about a plausible synthetic route: how many steps it would likely take, whether any step needs unusual or hazardous reagents, and how much stereochemical complexity needs to be controlled. 1 means the route looks impractical or unprecedented; 6 means a short, simple route from cheap starting materials.
Stability
How well the molecule would hold up in storage and handling as a bench chemical: its sensitivity to air, moisture, light, and heat, and whether it contains functional groups known to be reactive or short-lived on the shelf. 1 means the molecule would likely need special handling or wouldn’t keep at all; 6 means it should be as stable as a typical, well-behaved solid compound under normal storage.
This stability rating is specifically about physical shelf storage: it is not a prediction of how the molecule behaves metabolically inside a living organism, which is a separate question this rating doesn’t address.
Molecular Similarity Search
The Catalog page lets you search for molecules structurally similar to one you provide, rather than identical to it. Similarity is measured using Tanimoto similarity: a 0-to-1 score comparing the structural fingerprints of two molecules, where 1 means identical and lower values mean less structural overlap. This is the standard way computational chemistry tools compare “how alike” two structures are, without requiring an exact match.
Enter a SMILES string and a minimum similarity threshold, and the search returns every catalog molecule scoring at or above that threshold against your query.
Glossary
- SMILES
- Text notation for a molecule's structure. See Molecular Identity.
- SMARTS
- Text notation for a structural pattern, not one specific molecule. See Molecular Identity.
- InChI / InChI Key
- Standardized chemical identifier and its short hashed form. See Molecular Identity.
- MW
- Molecular Weight, the mass of a molecule. See Physicochemical Properties.
- LogP
- A measure of lipophilicity (fat- vs. water-solubility). See Physicochemical Properties.
- HBD / HBA
- H-Bond Donors / H-Bond Acceptors. See Physicochemical Properties.
- QED
- Quantitative Estimate of Drug-likeness, a 0-1 composite score. See Physicochemical Properties.
- SA Score
- Synthetic Accessibility Score, a computed 1-10 ease-of-synthesis estimate. See Physicochemical Properties.
- Log S
- Predicted aqueous solubility (log10 mol/L) from a trained graph neural network. See Physicochemical Properties.
- PSA
- Polar Surface Area. See Structural Characteristics.
- Rule of Five
- Lipinski's rule of thumb for oral bioavailability, based on MW/LogP/HBD/HBA. See Structural Characteristics.
- Stereocenter
- An atom whose spatial arrangement can produce distinct molecules with different activity. See Structural Characteristics.
- Target class
- A biological classification of drug targets (Kinase, GPCR, etc.). See Target Activity & Tiers.
- Antagonist / Agonist
- Blocking vs. activating a biological target. See Target Activity & Tiers.
- Tier (A-E)
- This platform's letter-grade summary of predicted target activity or safety, in place of a raw probability. See Target Activity & Tiers.
- CYP450
- A family of liver enzymes responsible for metabolizing many drugs. See Drug-Drug Interaction Screening.
- DDI
- Drug-Drug Interaction, the risk that a molecule interferes with how other drugs are metabolized. See Drug-Drug Interaction Screening.
- Tanimoto similarity
- A 0-1 score measuring structural similarity between two molecules' fingerprints. See Molecular Similarity Search.
- pLDDT
- Predicted Local Distance Difference Test — a structure prediction model's confidence in its own predicted 3D structure, 0-1. See Predicted Structure & Binding Affinity.
- Interface Confidence
- Confidence in a predicted structure's contact between two specific molecules, separate from overall pLDDT. See Predicted Structure & Binding Affinity.
- IC50
- The concentration of a molecule needed to inhibit a target's activity by 50% — a common measure of binding/inhibition strength. See Predicted Structure & Binding Affinity.
- RMSD
- Root-Mean-Square Deviation, a distance measure between two 3D poses of the same molecule. Used to compare Vina's independently-docked pose against Boltz's predicted pose. See Predicted Structure & Binding Affinity.