ResearchStrathclyde Text Analytics Group

Text analytics is concerned with inference from written communications. Modern developments in computational techniques and power have created opportunities to analyse vast quantities of text data to provide more effective decision support. The applications of these methods are widespread, the challenges are considerable but the potential benefits are substantial. As such, text analytics draws academics from across all four faculties of the university with interests ranging from developing fundamental techniques to applying methods to enhance outcomes.

The following provides an example of some of our work in Text Analytics.  If you would be interested in learning more about our activity either contact our group at textanalytics-group@strath.ac.uk or please make direct enquiries to particular academics.

Strathclyde Text Analytics Group - Seminar Series for 2026/27

Seminar Schedule 2026–2027
Date and Time Speaker Institution Title
Tuesday, October 6, 2026, 3pm Professor Adam Sanborn University of Warwick Understanding human cognition and perception as sampling
Tuesday, November 10, 2026, 3pm Professor Stephen Hansen University College London Policymakers’ Uncertainty
Tuesday, November 24, 2026, 3pm Nickil Maveli University of Edinburgh Coding and LLMs
Tuesday, January 19, 2027, 3pm Dr Kai-Robin Lange TU Dortmund Identifying economic narratives in large text corpora
Wednesday, February 10, 2027, 3pm Neha Singh University of Strathclyde Generative AI for long-term and complex decision-making
Thursday, March 11, 2027, 4pm Dr Margaret Leighton & Dr Irina Merkurieva University of St Andrews Childhood Aspirations and Adult Outcomes
Tuesday, April 20, 2027, 3pm Paolo Manildo University of Padua Advances in posterior computation of Latent Dirichlet Allocation models through informed non-reversible Markov chains
Tuesday, May 11, 2027,  3pm Dr Panagiotis Koutroumpis University of Reading Advanced textual analysis and language models for analysing board communications dynamics

Professor Adam Sanborn

  • Bio: Adam Sanborn is a Professor of Psychology at the University of Warwick. Professor Sanborn is interested in the rationality of human behaviour, which he studies with Bayesian models, approximations to Bayesian models, and behavioural experiments.
  • Format: Online
  • Title: Understanding human cognition and perception as sampling

Abstract

Over the past few decades, waves of complex probabilistic explanations have swept through cognitive science, explaining behaviour as tuned to environmental statistics in domains from intuitive physics and causal learning, to perception, motor control and language. Yet people produce stunningly incorrect answers in response to even the simplest questions about probabilities. How can a supposedly rational brain paradoxically reason so poorly with probabilities? Perhaps our minds do not represent or calculate probabilities at all and are, indeed, poorly adapted to do so. Instead, the brain could be approximating Bayesian inference through sampling: drawing samples from its distribution of likely hypotheses over time. Only with infinite samples does a Bayesian sampler conform to the laws of probability, and in this talk, I show how using a finite number of samples systematically generates classic probabilistic reasoning errors in individuals, and how an extended model explains estimates, choices, response times, and confidence judgments in a variety of tasks.

Professor Stephen Hansen

  • Bio: Stephen Hansen is a Professor of Economics at University College London and a Research Fellow at the Federal Reserve Bank of Dallas. He is also an Associate Editor for the Journal of Monetary Economics. Professor Hansen’s research uses unstructured data to build new measures of economic activity and behaviour across a variety of applications most often related to monetary policy and organisational economics.
  • Format: Online
  • Title: Policymakers’ Uncertainty

Abstract

Uncertainty is a ubiquitous concern emphasised by policymakers. We study how uncertainty affects decision-making by the Federal Open Market Committee (FOMC). We distinguish between the notion of Fed-managed uncertainty vis-a-vis uncertainty that emanates from within the economy and which the Fed takes as given. A simple theoretical framework illustrates how Fed-managed uncertainty introduces a wedge between the standard Taylor-type policy rule and the optimal decision. Using private Fed deliberations, we quantify the types of uncertainty the FOMC perceives and their effects on its policy stance. The FOMC's expressed inflation uncertainty strongly predicts a more hawkish policy stance that is not explained either by the Fed's macroeconomic forecasts or by public uncertainty proxies. We rationalize these results with a model of inflation tail risks and argue that the effect of uncertainty on the FOMC's decisions reflects policymakers' concern with maintaining credibility for the inflation anchor. 

Register

Nickil Maveli

  • Bio: Nickil Maveli is a PhD student at the University of Edinburgh’s Institute for Language, Cognition and Computation (ILCC), under supervision of Dr Shay Cohen.
  • Format: In-person
  • Title: To be confirmed

Abstract

To be confirmed.

Register

Dr Kai-Robin Lange

  • Bio: Dr Kai-Robin Lange is a postdoctoral researcher at the Department of Statistics, TU Dortmund University. His research interests encompass several topics in natural language processing, including content analysis of political debates, speeches and documents; tracing misinformation and conspiracy theories in social media; evaluation of embedding-based methods; extraction of events and narratives in text corpora; and spatio-temporal language modelling.
  • Format: Online
  • Title: Identifying economic narratives in large text corpora

Abstract

As economic and political narratives spread rapidly across digital media platforms, it has become increasingly critical to automatically extract such narratives from a corpus to gain an understanding of which narratives are spread by whom and on which platform. Previous pipelines attempting to extract narratives from a corpus of documents often employ a mix of state-of-the-art natural language processing techniques, such as BERT, to tackle this task. While effective on foundational linguistic operations essential for narrative extraction, such models lack the deeper semantic understanding required to distinguish extracting economic narratives from merely conducting classic tasks like Semantic Role Labeling.

Instead of relying on complex model pipelines, we evaluate the benefits of Large Language Models (LLMs) to extract narratives in one prompt using their great language understanding capabilities. We discuss two approaches: in a computationally heavy approach, we apply a rigorous narrative definition and compare GPT-4o interpretations of every single document in a corpus to gold-standard narratives produced by expert annotators. The second approach focuses on decreasing the computational demand of narrative analyses by pre-selecting documents of interest using a change-detection method based on the dynamic topic model RollingLDA. We discuss our findings and provide guidance for future work in economics and the social sciences that employs LLMs to pursue similar complex objectives.

Register

Neha Singh

  • Bio: Neha Singh is a PhD student in the Department of Management Science, University of Strathclyde, under supervision of Dr Euan Barlow.
  • Format: Online
  • Title: To be confirmed

Abstract

To be confirmed.

Register

Dr Margaret Leighton & Dr Irina Merkurieva

  • Bio: Dr Margaret Leighton is a Senior Lecturer in the Department of Economics, University of St Andrews. Her research is in applied microeconomics with a particular interest in education economics, as well as labour economics and development economics. Dr Irina Merkurieva is a Lecturer in the Department of Economics, University of St Andrews. Her research interests encompass labour economics, with particular interest in the dynamics of employment behaviour over the life cycle, search and matching, health and ageing.
  • Format: In-person
  • Title: Childhood Aspirations and Adult Outcomes

Abstract

This paper extracts aspirations from texts written in childhood by members of a British longitudinal cohort and explores how these relate to later life outcomes. Applying Natural Language Processing (NLP) tools to short essays collected at age 11, we identify four aspiration themes: family, hobbies, financial success and career. The weight of these four themes varies substantially across respondents, with girls on average placing more weight on family and boys on financial success.

Aspirations extracted using our method are strongly predictive of later life outcomes, even when controlling for detailed measures of early life environment, ability and family background. These associations are often highly heterogeneous by gender; for example, family-related aspirations are associated with higher educational attainment for men, but lower educational attainment for women.

Register

Paolo Manildo

  • Bio: Paolo Manildo is a PhD student in the Department of Statistics, University of Padua.
  • Format: Online
  • Title: Advances in posterior computation of Latent Dirichlet Allocation models through informed non-reversible Markov chains

Abstract

The Latent Dirichlet Allocation (LDA) is a probabilistic model which has become very popular in various scientific domains, for example natural language processing, where real data applications involve hundreds of documents, for a total of hundreds of thousands of words. Therefore, the standard collapsed Gibbs sampler, which updates the allocation of one word conditional on all others, often exhibits slow mixing and becomes infeasible as the total number of words grows.

Variational approximations have therefore been developed, which drastically reduce the computational burden. However, they can severely underestimate uncertainty and remain stuck in sub-optimal configurations. Leveraging recent results on non-reversible Markov chains for mixture models and informed proposals in discrete spaces, we introduce a novel sampling scheme for LDA designed to be efficient for large text corpora. We show both theoretically and empirically that this approach can significantly speed up the original algorithm with a modest increase in the cost per iteration.

Register

Dr Panagiotis Koutroumpis

  • Bio: Dr Panagiotis Koutroumpis is a Lecturer in Finance at the ICMA Centre, Henley Business School, University of Reading, whose research explores corporate finance, shadow banking, and the effects of geopolitical risk on financial markets.
  • Format: In-person
  • Title: To be confirmed

Abstract

To be confirmed.

Register

Current projects

James Bowden and Daniel Broby

We are using textual clues to identify the responsiveness of social reporting metrics to changes in sentiment. We use opinion mining and sentiment analysis to identify changes in how the public perceives the social responsiveness of a company. This will help us create a real time index that instantly captures when social metrics change. We hope it will provide a useful tool for the $15.02 trillion invested in funds using Socially Responsible screening.  Our contribution is expected to be the removal of significant subjective evaluation from the benchmarking of socially responsible investment.

Andrew Wodehouse, Jonathan Corney and Ross Maclachlan

Our research seeks to utilise the patent database more effectively in the engineering design process. In our latest work we have set out an approach that firstly utilises crowdsourcing to summarise patents and then applies text analysis in relation to three affective parameters: appearance, ease of use, and semantics. This has resulted in novel patent clusters that provide an alternative perspective on relevant technical data, and differs significantly from classifications using only functional requirements. The established interfaces and workflows emerging from the research support a new paradigm for the use of big data in engineering design, and could be applicable to other settings trying establish rich, user-centric information.

Zachary Greene

Systematically measuring policy goals has emerged as a major challenge to social scientists. Although traditional research tools such as surveys offer insights into priorities, these tools come with serious limitations and oftentimes reflect broader biases and the institutional context. Political scientists have turned to textual data from parliaments, interviews, newspaper and online sources to provide new insights. Politicians’ speeches reveal information over their policy preferences and the issues they care about. Likewise, the tone of debates predicts the broader mood towards an issue or politician. Scholars in the School of Government and Public Policy have used both supervised and unsupervised models for computational text analysis to help answer big political science questions. For example, work by Dr Greene uses speeches from party congresses to evaluate the disagreements between party members and the centrality of parties’ factions. Dr Brandenburg evaluates the tone of news coverage on the popularity of party leaders and major politicians. Other projects focus on social media networks and debates, the content of parties’ election programmes in diverse international settings and measuring newspapers’ bias towards candidates based on their gender and broader political background.

Full project title: Bridging the Research-Policy Gap in Entrepreneurial Ecosystems: A Computational Text Analysis Programme

This project employs computational text analysis to examine the extent to which academic research on entrepreneurial ecosystems is reflected in policy discourse across the four UK nations. By comparing thematic structures in academic literature and policy documentation across multiple case-study ecosystems, the project seeks to shed light on research utilisation gaps between scholars and policy-makers.

Researchers: Dr Cynthia Medeiros (Department of Management Science, University of Strathclyde), Dr Stephen Green (Edinburgh Business School, Heriot-Watt University)

Funding: This project has been funded by Heriot-Watt University's Early Career Development Programme.

Our members

Dr James Bowden

Co-lead

  • Accounting & Finance
  • University of Strathclyde

View James's staff profile

Dr Cynthia Medeiros

Co-lead

  • Management Science
  • University of Strathclyde

View Cynthia's staff profile