Jaemin Yang (Doctoral student in Informatics) has been working on ways to represent biomedical pathways using structures from nuclear energy. In the recent Computational Science Hackathon he used large language models to automate that process and placed second in the Ashby Prize. There is a great photo of Jaemin in the NCAS story (2025 Ashby Prize in Computational Science Hackathon Winners streamline CRISPR Screening with AI | Office of Data Science Research | Illinois) which ran the event. Well done Jaemin !
Category Archives: Uncategorized
Synergy Award
I am excited to share that the Midwest Big Data Innovation Hub (MBDH) won the Synergy Award from the Chicago Council on Science and Technology (C2ST). It was great to learn more about the Council’s inspiring efforts and an honor to accept the award on behalf of the many people who have contributed to the MBDH – particularly John MacMullen who has been serving as the Executive Director. The MBDH focused on how to bring people together to develop data-driven solutions that address challenges in the Midwest through seed grants, outreach, and community building. Although the project has now formally concluded the plethora of spin-off projects are an ongoing testament to how the combination of people and technology can address real challenges and I will personally value the connections that I made during my time on the MBDH forever.
The iSchool has a longer story with pictures Midwest Big Data Innovation Hub wins Synergy Award | School of Information Sciences
Innovation Grand Rounds: Using NLP to harness exposome factors that impact cancer
I am looking forward to giving a talk tomorrow (Fri Oct 14,12-1) at the MSB Auditorium (Room 274)
Healthy People 2030 and the All of Us Program demonstrate how human health is influenced by factors beyond direct patient/physician interactions, sometimes referred to as the exposome. A prime example is the increasing number of chemicals humans are exposed to and their effect on our overall chemical load. Commonly used risk assessment methods do not scale quickly or easily to changing conditions, such as localized regulatory changes. In this talk, Professor Catherine Blake will describe an automated approach that combines semantics and sentence syntax to identify target outcomes for cancer risk assessments from chemical exposure. In contrast to prior methods, this system would change a user’s decision for 21 out of the 27 chemicals explored. She will also discuss how this approach can be used to characterize other factors in the exposome that impact cancer.
Keynote: The case for Quality over Quantity
I will be presenting at the Health Science Librarians of Illinois (HSLI) Annual Conference on September 8 (https://hsli.org/conference/).
Evidence-based Librarianship and the Case for Quality over Quantity
Librarians have long played a critical role in teams that conduct systematic reviews. They partner with investigators to carefully construct search strategies that ensure all relevant studies are identified. To avoid publication bias, citations are hand searched and some teams even reach out to authors to identify “bottom drawer” literature, where a study has been conducted but not yet published. This example embodies the idea that getting the right data takes work but is critical to reach a valid conclusion. More importantly that easily accessible information can be misleading. This talk will provide a series of examples that illustrate when data quantity is no substitute for quality and how librarians are uniquely positioned within complex data ecosystems to ensure that the same biases that occur when using literature are not recreated or reinforced when researchers shift to re-using data.
automated risk assessment talk at George Washington University
I am looking forward to talking with folks from the Biomedical Informatics Center at George Washington University next week about “Using semantics to scale up evidence-based chemical risk-assessments and the implications to cancer research.”
Current methods used to conduct chemical risk assessments do not scale to recent regulatory changes such as the European Union’s REACH initiative that dramatically increases the number of chemicals to be assessed and the US EPAs trend towards cumulative risk assessments that consider multiple chemicals or combinations of chemical and non-chemical stressors. I will describe an automated approach that uses semantics to identify target outcomes and then captures supporting, refuting, or neutral with respect to the given target. The system scales to 482K abstracts and show that a user’s decision would change for 21 out of the 27 chemicals explored if only the retrieval step was considered instead of the extraction step. I will discuss the implications of this finding to cancer research.
This work is based on work with Jodi Flaws. More details can be found https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0260712
Using semantics to scale up evidence-based chemical risk-assessments
Just wanted to share our new paper, which uses explicit and observational claims from the Claim Framework to scale up evidence-based chemical risk-assessments, was just published https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0260712. This work was conducted with Jodi Flaws .
New paper at ASIS&T
Don Keefer presented our paper entitled The Reproducible Data Reuse (ReDaR) Framework to Capture and Assess Multiple Data Streams at ASIS&T last week (see the full paper here)
