Senior Data Scientist Biologics Discovery

Global healthcare company operating Innovative Medicine and MedTech businesses, including a robotic-assisted bronchoscopy platform.

New Brunswick, United States
About Johnson & Johnson

Johnson & Johnson develops medicines, therapies, medical devices and health technologies across oncology, immunology, neuroscience, cardiopulmonary, cardiovascular, orthopaedics, surgery and vision.

View jobs by Johnson & Johnson

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will analyze and integrate biologics discovery data, translate scientific questions into data requirements, and identify data-quality and modeling risks. You will develop features and model-ready datasets, define labels and aggregation with data engineers, and maintain traceable, reproducible datasets. You will collaborate with scientists, molecular-modeling, ontology, and MLOps colleagues to support responsible AI and reliable model use.

Requirements

  • Master’s or Ph.D. in Computer Science, Machine Learning, Computational Biology, Bioinformatics, Statistics, or a related field
  • At least 2 years of applied machine learning experience, including model development, evaluation, and dataset curation for complex scientific or biomedical data
  • Proficiency in Python, PyTorch, scikit-learn, and SQL
  • Experience converting heterogeneous experimental data into robust features and training sets
  • Exposure to cloud training and data infrastructure
  • Understanding of evaluation, validation, data leakage, and distribution shift
  • Ability to collaborate with experimental scientists and modeling partners in a matrixed R&D environment

Responsibilities

  • Analyze and integrate heterogeneous biologics discovery data
  • Translate scientific questions and design-make-test-learn decision points into data and analytical requirements
  • Identify data-quality issues, bias, leakage, and distribution-shift risks
  • Support molecule prioritization, risk identification, and hypothesis generation
  • Develop featurization approaches and model-ready datasets
  • Specify model features, labels, and aggregation with data engineers
  • Curate, document, and version reproducible datasets
  • Hand off standardized training datasets to molecular-modeling partners
  • Collaborate with ontology and MLOps colleagues
  • Champion reproducibility, documentation, and responsible AI

Benefits

  • Annual bonus or sales commissions, subject to pay grade and location
  • Vacation days
  • At least 12 weeks of parental leave
  • Bereavement leave
  • Caregiver leave
  • Volunteer leave
  • Well-being reimbursement
  • Financial, physical, and mental health programs
  • Service anniversary and recognition awards
  • Eligibility for insurance plans for employees and eligible dependents