The chembl-database-visualize skill provides programmatic query capabilities and visual analytics for EMBL-EBI's ChEMBL bioactivity database. It synthesizes compound properties, biological targets, concentration-normalized assays, and chemical similarity networks into self-contained, interactive web dashboards (dashboard.html).
The workflow diagram above illustrates the end-to-end architecture:
- Data Inputs: Ingests compound identifiers (ChEMBL IDs), biological target IDs, SMILES chemical notation, and bioactivity assay criteria.
- Skill Engine: Performs automated nanomolar (nM) unit normalization, Lipinski Rule of 5 physicochemical calculations, SAR bioactivity binning, and network graph modeling.
- Interactive Outputs and Dashboard: Emits standardized JSON/SDF data artifacts and compiles a standalone HTML/JavaScript dashboard featuring synchronized 2D chemical topologies, 3D WebGL rotating conformer viewports, and physicochemical radar charts.
- 3D Conformer and 2D Vector Structural Visualizations: Renders interactive WebGL conformers via 3Dmol.js alongside 2D vector diagrams with element-specific color coding.
- Physicochemical Profiling: Automated calculation of Lipinski Rule of 5 parameters (MW, AlogP, HBD, HBA, Rotatable Bonds, TPSA) with rule-violation detection and an SVG spider/radar chart.
- SAR and Bioactivity Distributions: Standardization of bioactivity measurements into nanomolar (nM) concentrations, computation of negative logarithmic affinity (pIC50), and rendering of potency histograms.
- Clinical Development Tracking: Stage timelines mapping compounds across Preclinical, Phase I, Phase II, Phase III, and Approved statuses, coupled with Mechanism of Action and ATC/MeSH indication cards.
- Chemical Similarity and Analog Matrix: Tanimoto similarity scoring and dynamic client-side filtering sliders for structure-activity exploration.
- Interactive Force-Directed Network Graph: Physics simulation mapping compound-target-assay connectivity with interactive node repositioning.
- Deterministic Offline Testing: Built-in
--mockengine executing synthetic test fixtures without external network dependencies. - Open-Source Licensing: Released under the Apache License 2.0 with automated compliance guards for upstream EMBL-EBI CC BY-SA terms.
.
|-- .gitignore
|-- .licenses/
| '-- chembl_database_visualize_LICENSE.txt
|-- assets/
| |-- chembl25_aspirin.json
| |-- chembl25_aspirin.sdf
| |-- chembl25_aspirin.svg
| |-- dashboard_schema.json
| |-- egfr_activities.json
| |-- similarity_aspirin.json
| |-- skill_workflow_dashboard.jpg
| '-- structure_fallback.svg
|-- dashboard.html
|-- documents/
| '-- prd.md
|-- LICENSE
|-- output/
| '-- aspirin/
| |-- activities_raw.json
| |-- dashboard.html
| |-- data.json
| |-- drug_summary.json
| |-- indications_raw.json
| |-- mechanisms_raw.json
| |-- molecule_raw.json
| |-- similarity_raw.json
| |-- structure.sdf
| '-- structure.svg
|-- pytest.ini
|-- README.md
|-- references/
| |-- api_endpoints.md
| |-- citation.bib
| '-- dashboard_schema.md
|-- requirements.txt
|-- sample_data/
| |-- chembl25_aspirin.json
| |-- chembl25_aspirin.sdf
| |-- chembl25_aspirin.svg
| |-- egfr_activities.json
| '-- similarity_aspirin.json
|-- scripts/
| |-- chembl_api.py
| |-- generate_dashboard.py
| |-- run_pipeline.py
| |-- visualize_bioactivity.py
| |-- visualize_compound.py
| '-- visualize_similarity.py
|-- SKILL.md
|-- SKILL_LICENSES.md
|-- skills/
| '-- chembl_database_visualize/
| |-- .licenses/
| | '-- chembl_database_visualize_LICENSE.txt
| |-- assets -> ../../assets
| |-- LICENSE -> ../../LICENSE
| |-- references -> ../../references
| |-- sample_data -> ../../sample_data
| |-- scripts -> ../../scripts
| '-- SKILL.md -> ../../SKILL.md
'-- tests/
|-- run_tests.py
|-- test_chembl_api.py
|-- test_pipeline.py
'-- test_visualizations.py
This skill acts as an automated research assistant for exploring medicinal chemistry and pharmacology data from the ChEMBL database. You do not need programming experience to understand and interact with its outputs.
- Drug Profiles: Look up any approved medicine or experimental molecule (such as Aspirin, Ibuprofen, or Imatinib) to inspect what conditions it treats, how it works in the body, and its 2D and 3D chemical structures.
- Drug-Likeness Rules: Evaluate whether a molecule has physical and chemical properties suitable for oral administration according to standard pharmaceutical criteria (Lipinski's Rule of 5), measuring molecular weight, water/fat solubility balance, and structural flexibility.
- Target Potency: Determine how strongly a molecule binds to a biological protein target (such as a receptor or enzyme). The skill converts experimental laboratory values from different studies into a unified scale (nanomolar concentrations and pIC50 scores).
- Related Compounds: Identify chemical analogs--molecules that share similar structural backbones--and compare their differences in size, solubility, and clinical development stage.
- Question to ask: "Tell me about the drug Aspirin (CHEMBL25). What medical conditions is it approved to treat, what enzymes does it block, and what does its structure look like?"
- What the skill does: The skill retrieves the clinical profile for Aspirin, identifies its approved indications (such as pain relief, fever reduction, and cardiovascular prophylaxis), extracts its mechanism of action (blocking Cyclooxygenase-1 and Cyclooxygenase-2 enzymes), verifies that it satisfies all four Lipinski oral drug-likeness criteria, and builds an interactive dashboard.
- How to explore the result:
Open the generated file
./output/aspirin/dashboard.html(or./dashboard.html) in any modern web browser. You can click and drag to rotate the 3D molecule, toggle functional group highlights on the 2D diagram, and switch between Light Mode and Dark Mode.
- Question to ask: "Summarize the bioactivity screening data for the receptor EGFR (CHEMBL203). How potent are the molecules tested against this target?"
- What the skill does:
The skill collects experimental inhibition measurements (
IC50values) across published screening studies, converts all concentration units to nanomolar (nM), calculates logarithmic affinity (pIC50), and groups the tested molecules into high, moderate, and weak potency categories. - How to explore the result: The generated dashboard displays a potency distribution histogram, a searchable table of laboratory assays, and a summary of the strongest binding compounds.
- Question to ask: "Find molecules with a chemical structure similar to Aspirin and compare their molecular weight and clinical stage."
- What the skill does: The skill runs a chemical fingerprint similarity search, ranking related molecules from 0% to 100% structural similarity (Tanimoto score), and compiles an analog comparison gallery.
- How to explore the result: Open the dashboard and use the interactive sliders to filter the analog cards in real time by minimum similarity percentage (for example, 85% or higher) and maximum molecular weight, or drag nodes in the force-directed network graph to inspect structural relationships.
- Python Version: Python 3.10 or higher.
- Operating System: Linux, macOS, or Windows.
- Environment Management: A virtual environment (
venv) oruvis recommended to isolate dependencies.
python3 -m venv .venv
source .venv/bin/activateInstall the required packages using pip:
pip install -r requirements.txtNote: The core scripts contain an in-tree fallback to Python's standard library urllib.request. If optional HTTP packages are omitted, requests execute using standard library utilities while preserving 5.0 QPS rate limiting and exponential backoff retry semantics.
The unified orchestrator script scripts/run_pipeline.py executes queries, downloads coordinates, processes bioactivity, and compiles the standalone interactive dashboard.
# 1. Full Compound Profiling Pipeline (e.g. Aspirin)
python3 scripts/run_pipeline.py --molecule CHEMBL25 -o ./output/aspirin/
# 2. Target SAR and Bioactivity Distribution (e.g. EGFR)
python3 scripts/run_pipeline.py --target CHEMBL203 --activity_type IC50 -o ./output/egfr/
# 3. Chemical Similarity and Analog Exploration
python3 scripts/run_pipeline.py --smiles "CC(=O)Oc1ccccc1C(=O)O" --similarity 85 -o ./output/analogs/
# 4. Offline Synthetic Execution (Sandboxed or CI environments)
python3 scripts/run_pipeline.py --molecule CHEMBL25 --mock -o ./output/mock_compound/The test suite includes smoke, unit, and end-to-end integration tests with synthetic data.
Execute tests via the standalone runner:
python3 tests/run_tests.pyOr execute tests via pytest:
pytest tests/Both test runners return standard exit codes (0 for success, non-zero for failure).
This section describes how an automated agent harness loads and interacts with this skill.
- Ensure the skill directory
skills/chembl_database_visualize/(or the repository root) is located in the agent harness skills directory (such asskills/). - The agent harness parses
SKILL.mdto register the available workflows, subcommands, and operational rules. - The license guard automatically verifies
.licenses/chembl_database_visualize_LICENSE.txtupon the first query, displaying licensing terms and logging an ISO 8601 UTC timestamp.
When integrating this skill into an AI agent system, the following prompt patterns guide the agent to invoke the appropriate workflows and deliver actionable outputs:
- User Prompt:
"Retrieve the chemical profile for Aspirin (CHEMBL25). Calculate Lipinski Rule of 5 drug-likeness metrics, fetch its 2D and 3D molecular structures, extract known mechanisms of action, and compile an interactive dashboard."
- Expected Agent Action:
Executes the unified pipeline with the compound identifier:
python3 scripts/run_pipeline.py --molecule CHEMBL25 -o ./output/aspirin/
- Delivered Output:
- Summarizes molecular formula (
C9H8O4), molecular weight (180.16 Da), AlogP (1.31), and 0 Lipinski violations. - Highlights primary mechanisms (Cyclooxygenase-1/2 inhibition) and approved indications.
- Provides a relative link to the interactive single-page dashboard at
./output/aspirin/dashboard.html.
- Summarizes molecular formula (
- User Prompt:
"Query the bioactivity profile for the target EGFR (CHEMBL203). Standardize all reported IC50 values to nanomolar units, compute pIC50 values, categorize potency distributions, and generate a visual analytics dashboard."
- Expected Agent Action:
Executes the pipeline targeting the specified protein:
python3 scripts/run_pipeline.py --target CHEMBL203 --activity_type IC50 -o ./output/egfr/
- Delivered Output:
- Reports total assay count, standard concentration ranges, and potency counts across High (<100 nM), Moderate (100 to 10,000 nM), and Weak (>10,000 nM) bins.
- Delivers
./output/egfr/dashboard.htmlcontaining the potency histogram and search-filterable assay table.
- User Prompt:
"Run a chemical similarity search for the molecule with SMILES 'CC(=O)Oc1ccccc1C(=O)O' at an 85% Tanimoto threshold. Find structurally related analogs and generate an analog comparison grid with interactive sliders."
- Expected Agent Action:
Executes the similarity search workflow:
python3 scripts/run_pipeline.py --smiles "CC(=O)Oc1ccccc1C(=O)O" --similarity 85 -o ./output/analogs/ - Delivered Output:
- Lists top analog hits (such as Salicylic Acid, Diflunisal, Salsalate) with similarity scores and molecular weights.
- Delivers
./output/analogs/dashboard.htmlfeaturing interactive client-side sliders for similarity and molecular weight thresholding.
- User Prompt:
"Look up the mechanisms of action and reported clinical indications for CHEMBL25. Which target proteins does it interact with, and what clinical phases has it reached?"
- Expected Agent Action:
Queries specific entity endpoints using the modular API utility:
python3 scripts/chembl_api.py mechanism --filter molecule_chembl_id=CHEMBL25 --output ./output/mechanisms.json python3 scripts/chembl_api.py drug_indication --filter molecule_chembl_id=CHEMBL25 --output ./output/indications.json
- Delivered Output:
- Provides structured JSON exports and textual synthesis of mechanisms (direct interaction with COX enzymes) and indications (analgesic, antipyretic, anti-inflammatory, antiplatelet) with max clinical phase 4.
- User Prompt:
"Validate the visualization pipeline offline for CHEMBL25 without making live network requests. Verify that all dashboard and data artifacts are created correctly."
- Expected Agent Action:
Executes the pipeline with the
--mockflag enabled:python3 scripts/run_pipeline.py --molecule CHEMBL25 --mock -o ./output/mock_aspirin/
- Delivered Output:
- Confirms generation of
./output/mock_aspirin/dashboard.html,./output/mock_aspirin/data.json,./output/mock_aspirin/structure.sdf, and./output/mock_aspirin/structure.svgusing synthetic sample fixtures.
- Confirms generation of
The individual components in scripts/ can also be invoked independently:
scripts/chembl_api.py: Low-level rate-limited REST client for all ChEMBL API entities.scripts/visualize_compound.py: Physicochemical property calculations and Lipinski scoring.scripts/visualize_bioactivity.py: Concentration normalization (nM), pIC50 computation, and potency binning.scripts/visualize_similarity.py: Tanimoto similarity parsing and analog scoring.scripts/generate_dashboard.py: Single-page interactive HTML dashboard compiler.
When utilizing data generated by this skill, cite the following references:
- ChEMBL 2024 Release: Zdrazil, B. et al. The ChEMBL Database in 2023: a drug discovery platform spanning multiple bioactivity data types and time periods. Nucleic Acids Research 52, D1180-D1192 (2024). doi:10.1093/nar/gkad1004.
- ChEMBL Web Services: Davies, M. et al. ChEMBL web services: streamlining access to drug discovery data and utilities. Nucleic Acids Research 43, W612-W620 (2015). doi:10.1093/nar/gkv352.
ChEMBL bioactivity values, molecular properties, and predicted data are aggregated from scientific literature and high-throughput screening assays. They are intended for research and informational purposes only and must be experimentally validated before medicinal, diagnostic, or clinical application.
This project is licensed under the Apache License 2.0. A copy of the full license text is provided in the LICENSE file located in the project root directory.
- Source Code and Documentation: All Python scripts in
scripts/, test suites intests/, documentation files (README.md,SKILL.md,documents/prd.md), and configuration schemas inassets/andreferences/are released under the terms of the Apache License 2.0. - Database Content Attribution: Scientific data extracted from the European Bioinformatics Institute (EMBL-EBI) ChEMBL database is redistributed under Creative Commons Attribution-ShareAlike (CC BY-SA 3.0 / CC BY-SA 4.0). Users must review and adhere to upstream ChEMBL licensing terms when utilizing or publishing datasets derived from this skill.
- License Registry: The licensing terms and upstream references are cataloged in
SKILL_LICENSES.mdand monitored via.licenses/chembl_database_visualize_LICENSE.txt.
