About the role#
This co-op position supports the Advanced Clinical Research group in Danvers. You will develop methods to extract clinical data from large collections of source documents, focusing on transforming PDF-based records into structured databases. This work supports clinical research, real-world evidence generation, and scientific analysis. The role combines clinical research, data science, artificial intelligence, and software development, working alongside clinical scientists to validate automated data extraction workflows.
What you'll do#
- Develop methods for extracting structured data from large collections of PDF-based clinical source documents.
- Assist in evaluating and implementing document-processing technologies, including OCR, NLP, and AI-assisted extraction.
- Design workflows to convert unstructured clinical information into standardized datasets.
- Develop and maintain databases for storing extracted clinical data.
- Create and optimize SQL queries to support data quality review and downstream research.
- Validate extracted data against source documents and reference datasets to assess accuracy and completeness.
- Perform quality control reviews and document extraction errors, inconsistencies, and improvement opportunities.
- Harmonize extracted variables according to predefined data dictionaries and research specifications.
- Support the development of automated pipelines that improve the efficiency and scalability of clinical data collection.
- Contribute to multiple clinical research initiatives across the Advanced Clinical Research team.
What you'll need#
- Currently enrolled in a Bachelor's or Master's degree program in Computer Science, Data Science, Biomedical Informatics, Bioinformatics, Biomedical Engineering, Software Engineering, Health Informatics, Statistics, Applied Mathematics, or a related quantitative discipline.
- Candidates from health sciences backgrounds with strong programming experience are encouraged to apply.
- Experience programming in Python.
- Working knowledge of SQL and relational databases.
- Strong analytical and problem-solving skills.
- Ability to work with large datasets and perform data cleaning and transformation tasks.
- Strong attention to detail and commitment to data quality.
- Ability to follow structured validation and documentation procedures.
- Excellent written and verbal communication skills.
- Ability to work independently while collaborating within a multidisciplinary research team.
- Preferred: Experience processing PDFs, text files, or other unstructured data sources.
- Preferred: Familiarity with OCR technologies, document-processing tools, machine learning, NLP, or LLMs.
- Preferred: Experience building data pipelines, automated workflows, or using Git.
- Preferred: Familiarity with PostgreSQL, SQL Server, MySQL, or similar database platforms.
- Preferred: Exposure to healthcare, clinical, or biomedical datasets and coursework involving artificial intelligence or data engineering.
Location & details#
- Location: Danvers, Massachusetts.
- Modality: On-site.
- Employment: Full-time.
- Term: Rolling.
- Note: This position does not provide visa sponsorship.
About Johnson & Johnson
Johnson & Johnson is a public company based in New Brunswick, New Jersey. Founded in 1886, it operates within the hospitals and health care industry. The company focuses on medical devices, diagnostics, and pharmaceuticals. It employs over 133,000 people across its global offices.
How to get in at Johnson & Johnson
Applying early is a significant advantage when pursuing a role at Johnson & Johnson, as recruiters often review candidates before the applicant pool becomes unmanageable. Intern Insider sends an instant alert the moment a role matching your target is published, helping you get your application in among the first. It is a simple way to stay ahead of the crowd during a busy recruiting season. You can also improve your response rates by reaching out to recruiters directly to ask about the role or a potential referral. Intern Insider surfaces the specific recruiters behind the company's roles, which allows you to bypass the general queue and connect with the people actually making the hiring decisions.



