Customer case
GenAI in Pharma: Molecule Discovery Automation with NNIT
Analysts from a leading pharmaceutical company spent significant time manually gathering molecule data from regulatory and commercial sources. NNIT built an automated pipeline that delivers a structured, ready-to-evaluate shortlist instead.
GenAI in Pharma R&D: From Manual Search to Automated Shortlist
A global leading pharmaceutical company worked with NNIT's data and AI services for life sciences to replace the labour-intensive task of manually gathering molecule data with an automated pipeline that identifies generic drug development opportunities. The pipeline analysed over 10,000 molecules during the PoC, delivering a dataset of approximately 1,500 new candidates to the company.
Built on Azure Databricks, the proof-of-concept pipeline extracts, standardises and consolidates molecule data from six external sources into a single structured output aligned with the company’s internal evaluation schema. During the proof of concept, it analysed over 10,000 molecules and delivered a shortlist of 1,500 candidates, so analysts receive a pre-populated starting point for their commercial opportunity assessments rather than a blank sheet.
Case in brief:
Challenge: Analysts manually collected and cross-referenced molecule data from multiple regulatory and commercial sources to find generic drug opportunities.
Solution: An AI-powered pharmaceutical data pipeline on Azure Databricks that extracts, standardises and consolidates data from six external sources.
Benefit: Over 10,000 molecules analysed in a single pipeline run, producing a shortlist of 1,681 candidates enriched with regulatory status, market size estimates and formulation details.
Manual Data Gathering Slowed Generic Drug Development
The company is a global generic pharmaceutical company. Continuously scanning the public molecule data landscape for new opportunities is therefore a key part of its R&D workflow. That process normally requires multiple analysts to manually search clinical trial registries, regulatory databases, and commercial intelligence platforms, then cross-reference and consolidate the findings in their internal evaluation tool.
The work is time-intensive, prone to inconsistencies and difficult to scale, particularly as the number of relevant molecules in late-stage clinical development continues to grow. The company wanted to automate the initial data intake and give analysts a pre-populated, structured starting point for the commercial opportunity assessments in their early R&D pipeline
Our customer previously relied on analysts to manually gather and cross-reference molecule data from multiple regulatory and commercial sources. With NNIT’s GenAI-powered data pipeline on Azure Databricks, we automate ingestion, standardisation and consolidation from six external sources – turning a landscape of 10,000+ molecules into a structured, evaluation-ready shortlist of 1,500 candidates in a single run.
Joachim Breitenstein, Business Consultant, NNIT
A Pharmaceutical Data Pipeline Built on Databricks for Life Sciences
NNIT designed and delivered a proof-of-concept data pipeline running on Azure Databricks that automatically ingests data from six external sources and consolidates it into a unified molecule-level dataset matching the company’s internal evaluation schema. The pipeline runs on NNIT's cloud and data engineering infrastructure and works in three stages. The extraction stage connects to each source and downloads the latest available data. The transformation stage cleans and standardises what was collected into a consistent format.
The consolidation stage merges everything into a single structured molecule table, covering over 10,000 molecules per run. The pipeline is configured from key deterministic business logic already used at the company, which provides the rules for mapping and filtering the data. GenAI is then used to estimate the more complex, commercial data points where no structured source provides them directly.
Insights
Life Sciences, AI, Clinical, Data
Data governance for life sciences - A fast-growing biotech company gained a data foundation built to scale with its growth
Life Sciences, Data
Master Data Management for Faster, More Compliant Pharma Decisions
Life Sciences, Clinical, Compliance, Digital Manufacturing, Quality, Regulatory Affairs, Smart supply-chain
Webinar: Inside the NDC-12 Industry Pilot
Life Sciences, Clinical, Regulatory Affairs
Webinar: Content Governance in the Age of AI: Preparing Pharmaceutical Organisations for ePI and FHIR I by Docuvera & NNIT
Life Sciences, Data, Digital Manufacturing
Webinar: Advancing Digital Manufacturing with Ignition 8.3
Life Sciences, Data, Digital Manufacturing
Webinar: Unified Namespace: Building a Scalable Data Hub for Manufacturing
Life Sciences, AI
Webinar: Experience AI that Solves your Life Sciences Challenges
Life Sciences, AI, Digital Manufacturing
Accelerating PAS-X Upgrades with Structured AI Automation
Life Sciences, Data, Regulatory Affairs, Smart supply-chain
The NDC-12 Cross-Industry Pilot: Key Findings & Insights From Phase 0
Pharma Domain Knowledge Applied to AI in Drug Discovery
NNIT combined pharmaceutical domain knowledge with data engineering and AI expertise to bridge the gap between raw public data and the company's internal evaluation workflow. The NNIT teams designed the solution to mirror how the analysts themselves would prioritise and map the data. The result is a technical data pipeline and foundation designed to accelerate the molecule discovery workflow for pharmaceutical generics development, turning a landscape of over 10,000 molecules into 1,500 evaluation-ready candidates, all produced automatically from a single pipeline run.
Alera was leveraged throughout the project for code development and documentation writing to accelerate delivery while maintaining quality.