Carolyn Penstein Rosé profile photo

Carolyn Penstein Rosé

Professor

  • Pittsburgh PA UNITED STATES

Carolyn Rosé's research advances Sociotechnical AI through human-AI collaboration and data-driven inductive biases and representations.

Contact

Biography

Dr. Carolyn Rosé is the Kavčić-Moura Professor of Language Technologies and Human-Computer Interaction at Carnegie Mellon University's School of Computer Science. She holds a Masters in Computational Linguistics (1994) and PhD in Language and Information Technologies (1998), both from CMU, and has been CMU faculty since 2003.

Her university leadership includes serving as Interim Director of the Language Technologies Institute, Director of the Master of Computational Data Science program, Director of Language Technologies Undergraduate Programs, Faculty Senate member, and Vice Chair of the University Education Council. She launched two highly ranked online certificates in Machine Learning/Data Science and Generative AI, and directed CMU's Generative AI Innovation Incubator in 2023. She has mentored 23 PhD students, over 100 masters students, 9 postdoctoral fellows, and 20+ undergraduates.

Externally, she has held board, editorial, and conference leadership roles across major research societies. In the ACL community, she served as Program Co-Chair for EMNLP 2025, keynote speaker at ACL 2018, and holds editorial and advisory roles in journals and committees. She is a Past President and Inaugural Fellow of the International Society of the Learning Sciences, Founding Chair of the International Alliance to Advance Learning in the Digital Era, Senior IEEE Member, and Executive Editor of the International Journal of Computer-Supported Collaborative Learning (impact factor 6.8). She is also a AAAS Leshner Leadership Institute Fellow (AI Cohort, 2020–2021).

Industry partnerships have spanned Google, Microsoft, Amazon, Oracle, Abridge AI, the Gates Foundation, and the Schmidt Foundation, among others. At the government level, she collaborated for six years with NLP teams at the NIH and Social Security Administration on decision support technology.

Areas of Expertise

Human-AI Collaboration
Computer-Supported Collaborative Learning
Conversational Agents
Automated Discourse Analysis
Medical NLP
Coding Agents
Computational Data Science
Computational Social Sciences

Media Appearances

Finalist teams advance in the Amazon Nova AI Challenge: Trusted AI Track

Amazon Science  online

2025-06-24

Since November 2024, ten top university teams from around the world have competed in the inaugural Amazon Nova AI Challenge: Trusted AI Track, focused on strengthening security in AI coding assistants and developing new automated methods to ‘red-team’ and test them. After months of intense competition, eight teams have advanced to the finals, demonstrating outstanding innovation in securing AI-powered code generation.

View More

The AI company Elon Musk cofounded just released a 'groundbreaking' tool that can automatically mimic human writing — here's how stunned developers are experimenting with it so far

Business Insider  online

2020-07-22

"Historically, natural language generation systems have lacked some nuance," said Carolyn Rose, a professor at Carnegie Mellon University's Language Technologies Institute. But GPT-3 seems different. Based on early reactions, GPT-3 has blown past existing models thanks to its massive dataset and its use of 175 billion parameters — rules the algorithm relies on to decide which word should come next to mimic conversational English. By comparison, the previous version, GPT-2, utilized 1.5 billion parameters, and the next most powerful model — from Microsoft — has 17 billion parameters.

View More

Elon Musk-Backed AI Company Launches New Tool that Writes Naturally Like Humans

Tech Times  

2020-07-22

Carnegie Mellon University's Language Technologies Institute Professor Carolyn Rose told Business Insider that while natural language generation systems have historically "lacked some nuance," GPT-3 seems different.

View More

Media

Social

Industry Expertise

Education/Learning
Computer Software

Education

Carnegie Mellon University:

Ph.D.

Language and Information Technologies

1997

Carnegie Mellon University

M.S.

Computational Linguistics

1994

University of California at Irvine

B.S.

Information and Computer Science

1992

Articles

Adapting to the Long Tail: A Meta-Analysis of Transfer Learning Research for Language Understanding Tasks

Transactions of the Association for Computational Linguistics

2022

Natural language understanding (NLU) has made massive progress driven by large benchmarks, but benchmarks often leave a long tail of infrequent phenomena underrepresented. We reflect on the question: Have transfer learning methods sufficiently addressed the poor performance of benchmark-trained models on the long tail? We conceptualize the long tail using macro-level dimensions (underrepresented genres, topics, etc.), and perform a qualitative meta-analysis of 100 representative papers on transfer learning research for NLU. Our analysis asks three questions: (i) Which long tail dimensions do transfer learning studies target? (ii) Which properties of adaptation methods help improve performance on the long tail? (iii) Which methodological gaps have greatest negative impact on long tail performance?

View more

Examining socially shared regulation and shared physiological arousal events with multimodal learning analytics

British Journal of Educational Technology

2023

Socially shared regulation contributes to the success of collaborative learning. However, the assessment of socially shared regulation of learning (SSRL) faces several challenges in the effort to increase the understanding of collaborative learning and support outcomes due to the unobservability of the related cognitive and emotional processes. The recent development of trace‐based assessment has enabled innovative opportunities to overcome the problem. Despite the potential of a trace‐based approach to study SSRL, there remains a paucity of evidence on how trace‐based evidence could be captured and utilised to assess and promote SSRL. This study aims to investigate the assessment of electrodermal activities (EDA) data to understand and support SSRL in collaborative learning, hence enhancing learning outcomes

View more

High school students’ data modeling practices and processes: From modeling unstructured data to evaluating automated decisions

Learning, Media and Technology

2023

It’s critical to foster artificial intelligence (AI) literacy for high school students, the first generation to grow up surrounded by AI, to understand working mechanism of data-driven AI technologies and critically evaluate automated decisions from predictive models. While efforts have been made to engage youth in understanding AI through developing machine learning models, few provided in-depth insights into the nuanced learning processes. In this study, we examined high school students’ data modeling practices and processes. Twenty-eight students developed machine learning models with text data for classifying negative and positive reviews of ice cream stores. We identified nine data modeling practices that describe students’ processes of model exploration, development, and testing and two themes about evaluating automated decisions from data technologies.

View more