sep
English language and linguistics research seminar: Esme Richarson-Owen, Salzburg University: Recruiting an LLM for large scale linguistic annotation tasks: assessing consistency and reliability.
Manual linguistic annotation is labour-intensive, time-intensive and prone to inter-annotator disagreement, all of which limit the scale at which patterns emerging from linguistic data requiring qualitative judgements can be generalised. Large Language Models (LLMs) offer a potential alternative, but their reliability across different annotation tasks remains underexplored. This study tests an LLM (DeepSeek-v3.2) on sensory adjective-noun constructions sourced from the British National Corpus 2014. The study serves as a pilot for my larger doctoral project that aims to scale this pipeline further.
Using a CRISPE-structured prompt, the LLM annotated c. 2,000 concordance lines for sentiment, evaluation and semantic classification across five iterations at two different temperature settings, with outputs compared against human annotation of the same data. Reliability was measured using Krippendorff’s Alpha (intra-annotator reliability) and Cohen’s Kappa (LLM-human agreement).
The findings address whether LLMs can be reliably recruited for large-scale, but fine-grained, linguistic annotation tasks, discussing some of the benefits to the approach, along with some of the potential drawbacks along with issues concerning reproducibility and replicability when LLMs are used for this type of task.
Om händelsen:
Plats: H339
Kontakt: panos.athanasopoulosenglund.luse
