All research topics

Research topic

Large Language Models

Large neural language models that learn broad linguistic patterns and can be adapted for many text tasks.

Overview

Large language models are trained on extensive text collections to predict and model sequences of tokens. Their capabilities emerge from scale, architecture, data mixture, optimization, and post-training, while their limitations include unreliable factuality, bias, and sensitivity to context.

What it is

An LLM is a parameterized model that assigns probabilities to token sequences and uses those probabilities to generate or score text. Post-training can make it more useful for dialogue, instruction following, coding, or specialized domains.

How it works

Text is tokenized, transformed into vectors, and processed through repeated layers, commonly using self-attention. Pretraining learns statistical regularities; supervised examples and preference-based methods can shape behavior afterward. Inference then generates tokens step by step under a decoding strategy.

Key concepts

  • Tokenization and embeddings
  • Pretraining and scaling laws
  • Instruction tuning
  • Alignment and preference optimization
  • Context windows
  • Factuality and calibration

Current research questions

  • What explains sudden or uneven capability changes with scale?
  • How can models cite evidence and distinguish knowledge from uncertainty?
  • How should data provenance, memorization, and privacy be evaluated?
  • Which evaluations measure useful reasoning rather than benchmark familiarity?

Applications

  • Writing and coding assistance
  • Document classification and extraction
  • Tutoring and language learning
  • Research support
  • Conversational interfaces

Relevant research papers

A live selection from the ScholarLens research index. Open any result to read its full paper page and use the Reading Assistant where available.

Finding relevant papers...

Keep exploring

Search the literature with your own question, or open a paper and read it closely with ScholarLens.