In my last article, I shared my unexpected dive into protein research, starting with databases like the Electron Microscopy Data Bank (EMDB) and the Protein Data Bank (PDB). What started as a curiosity quickly became a deep dive into how artificial intelligence (AI) can be trained to understand protein structures—something I never imagined I’d be exploring without a PhD.
This time, I’m going even deeper—but not in a way that requires years of specialized training. Instead, I want to highlight a powerful, often-overlooked entry point into cancer research: data annotation. Specifically, I’m focusing on pancreatic cancer and how everyday people (with the right tools) can contribute by labeling and organizing biological data.
Proteins play a crucial role in health and disease, including conditions like pancreatic cancer. But simply having raw protein data isn’t enough—without clear labels and classifications, this data is just a collection of numbers and structures with no context. That’s where annotation comes in.
Annotation involves adding meaningful labels to protein structures, mutations, and interactions, making it possible for both humans and AI to recognize patterns. Platforms like the Electron Microscopy Data Bank (EMDB) and Protein Data Bank (PDB) store this structured data, allowing researchers to study how proteins misfold or interact with drugs.
Without standardized annotations, AI models can’t make accurate predictions—just like a student trying to study from a textbook with missing labels and scrambled chapters. Good annotation turns raw data into knowledge, helping AI uncover patterns that could lead to better cancer treatments.
One of the biggest misconceptions about AI in science is that it replaces human expertise. In reality, AI is only as good as the data it learns from—and that data needs careful labeling, classification, and verification by humans before an AI model can recognize patterns.
I stumbled upon this firsthand while exploring BioPython, a tool that connects to databases like NCBI’s Entrez and the European Bioinformatics Institute’s EMBL-EBI. Through BioPython, I saw how annotation frameworks like Gene Ontology (GO) and Sequence Ontology (SO) provide structured ways to define protein functions, mutations, and interactions.
While this might sound technical, it’s surprisingly approachable. Just as you don’t need to be a software engineer to teach an AI to recognize handwritten digits, you don’t need a PhD to help AI recognize biologically relevant patterns.
For those wondering “Where do I start?”, here are some simple ways to explore protein annotation without a formal science background:
Try BioPython – This open-source tool allows you to retrieve and analyze protein sequences. Think of it as a bridge between complex biological data and accessible programming tools. (Biopython Docs)
Explore Open Datasets – Databases like EMDB (Electron Microscopy Data Bank) and PDB (Protein Data Bank) offer thousands of protein structures available for annotation and research.
Participate in Open Science Challenges – Platforms like Kaggle host competitions where you can contribute to projects like 3D particle annotation in cryo-electron tomography (CryoET)—even with minimal prior experience.
Learn About Gene and Sequence Ontology – These frameworks define biological features in a structured way, making it easier for AI models to learn from human-labeled data.
One of the most impactful contributions to pancreatic cancer research isn’t just in discovering new data—it’s in organizing and annotating existing data so it becomes usable. In CryoET annotation, where 3D imaging reveals protein structures at the molecular level, properly labeled data helps AI models and researchers quickly identify protein misfolding, structural abnormalities, and drug interactions. By focusing on cleaning, structuring, and annotating CryoET datasets, we ensure that when it’s time to analyze the data, scientists and AI tools can efficiently extract insights, accelerating discoveries that could lead to better diagnostics and treatments.
Above, search result from The Electron Microscopy Data Bank.
The world of protein research is no longer reserved for academic institutions or biotech labs. Thanks to open data initiatives and AI-powered tools, anyone with curiosity and a willingness to learn can contribute to groundbreaking discoveries.
By starting with something as simple as annotating protein datasets, we can collectively build better AI models that may one day help uncover new treatments for diseases like pancreatic cancer. And if I, a non-scientist, can navigate this space—so can you.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.