nnit_customer_case_logo_no_logo.png

Customer case

GenAI in Pharma: Molecule Discovery Automation with NNIT

Analysts from a leading pharmaceutical company spent significant time manually gathering molecule data from regulatory and commercial sources. NNIT built an automated pipeline that delivers a structured, ready-to-evaluate shortlist instead.

NNIT colleagues working together in a modern office focused on regulated industries.

GenAI in Pharma R&D: From Manual Search to Automated Shortlist

A global leading pharmaceutical company worked with NNIT's data and AI services for life sciences to replace the labour-intensive task of manually gathering molecule data with an automated pipeline that identifies generic drug development opportunities. The pipeline analysed over 10,000 molecules during the PoC, delivering a dataset of approximately 1,500 new candidates to the company.

Built on Azure Databricks, the proof-of-concept pipeline extracts, standardises and consolidates molecule data from six external sources into a single structured output aligned with the company’s internal evaluation schema. During the proof of concept, it analysed over 10,000 molecules and delivered a shortlist of 1,500 candidates, so analysts receive a pre-populated starting point for their commercial opportunity assessments rather than a blank sheet.

Case in brief:

Challenge: Analysts manually collected and cross-referenced molecule data from multiple regulatory and commercial sources to find generic drug opportunities.

Solution: An AI-powered pharmaceutical data pipeline on Azure Databricks that extracts, standardises and consolidates data from six external sources.

Benefit: Over 10,000 molecules analysed in a single pipeline run, producing a shortlist of 1,681 candidates enriched with regulatory status, market size estimates and formulation details.

Manual Data Gathering Slowed Generic Drug Development

The company is a global generic pharmaceutical company. Continuously scanning the public molecule data landscape for new opportunities is therefore a key part of its R&D workflow. That process normally requires multiple analysts to manually search clinical trial registries, regulatory databases, and commercial intelligence platforms, then cross-reference and consolidate the findings in their internal evaluation tool.

The work is time-intensive, prone to inconsistencies and difficult to scale, particularly as the number of relevant molecules in late-stage clinical development continues to grow. The company wanted to automate the initial data intake and give analysts a pre-populated, structured starting point for the commercial opportunity assessments in their early R&D pipeline

Our customer previously relied on analysts to manually gather and cross-reference molecule data from multiple regulatory and commercial sources. With NNIT’s GenAI-powered data pipeline on Azure Databricks, we automate ingestion, standardisation and consolidation from six external sources – turning a landscape of 10,000+ molecules into a structured, evaluation-ready shortlist of 1,500 candidates in a single run.

Joachim Breitenstein, Business Consultant, NNIT

A Pharmaceutical Data Pipeline Built on Databricks for Life Sciences

NNIT designed and delivered a proof-of-concept data pipeline running on Azure Databricks that automatically ingests data from six external sources and consolidates it into a unified molecule-level dataset matching the company’s internal evaluation schema. The pipeline runs on NNIT's cloud and data engineering infrastructure and works in three stages. The extraction stage connects to each source and downloads the latest available data. The transformation stage cleans and standardises what was collected into a consistent format.

The consolidation stage merges everything into a single structured molecule table, covering over 10,000 molecules per run. The pipeline is configured from key deterministic business logic already used at the company, which provides the rules for mapping and filtering the data. GenAI is then used to estimate the more complex, commercial data points where no structured source provides them directly.

Pharma Domain Knowledge Applied to AI in Drug Discovery

NNIT combined pharmaceutical domain knowledge with data engineering and AI expertise to bridge the gap between raw public data and the company's internal evaluation workflow. The NNIT teams designed the solution to mirror how the analysts themselves would prioritise and map the data. The result is a technical data pipeline and foundation designed to accelerate the molecule discovery workflow for pharmaceutical generics development, turning a landscape of over 10,000 molecules into 1,500 evaluation-ready candidates, all produced automatically from a single pipeline run.

Alera was leveraged throughout the project for code development and documentation writing to accelerate delivery while maintaining quality.

Portrait of an IT consultant who is a expert within the field of AI.

Talk to our AI Experts

Let’s scope your roadmap, de risk your compliance, and unlock rapid ROI from AI

When you submit your inquiry to NNIT via the contact form, NNIT process the collected personal data in accordance with the Privacy Notice, where you can read more about your rights and how NNIT process your personal data.