Skip to content
Mvaak

04 β€” Data Type

Text Annotated Data

Structured language datasets for NLP, LLM and conversational AI.

What it is

It looks simple until two people disagree.

Text annotation looks simple until two annotators disagree about the same sentence β€” and then it becomes a question of definitions.

Most of the value in a language dataset comes from resolving those definitions up front: what an entity boundary includes, how ambiguity is handled, and what to do with examples that genuinely sit between two labels.

Mvaak treats the guideline as the deliverable's foundation, and measures agreement against it rather than assuming it.

Why it matters for AI

Three things that decide whether it trains well.

  1. 01

    Disagreement is information

    When two annotators split on the same example, the label set is usually underspecified. Treating that as a signal rather than an error is what stops the ambiguity reaching the model.

  2. 02

    Boundaries are a decision, not a fact

    Whether a title, a qualifier or a trailing clause belongs inside an entity span changes what the model learns. Left to individual judgement, it changes per annotator.

  3. 03

    The awkward middle carries the weight

    Clear examples teach a model little it could not infer. The examples that sit between two labels are where a dataset earns its cost β€” and where a guideline has to be explicit.

Applications

What this data trains.

Each of these draws on the category differently, which is why the specification is written per project rather than per service.

  • NLP
  • LLM Training
  • Text Classification
  • Entity Recognition
  • Intent Classification
  • Sentiment Analysis
  • Conversational AI
  • Language Intelligence

Collection possibilities

Where the data comes from.

Most engagements are a mix: annotate what you already hold, and collect specifically where coverage is thin.

Annotating your corpus

Most text work starts from documents, transcripts, tickets or conversations you already hold, labelled against a schema defined with your team.

Targeted collection

Where a class is rare, examples can be sourced or written specifically to cover it, so the label is not learned from a handful of instances.

Guideline development

Where the definitions are not settled, the first pass is a small annotated sample used to expose disagreement and tighten the schema before volume begins.

Annotation possibilities

Scoped to the project, not sold as a bundle.

These are the operations available for this data type. Which of them a project uses depends on what the model has to learn.

  • 01Classification and tagging
  • 02Named entity annotation
  • 03Intent labelling
  • 04Sentiment annotation
  • 05Span and relation marking
  • 06Human validation of edge cases

Workflow

The same five stages, every project.

01

Understand

We start from the model and the use case, not the file format β€” what the system needs to learn, and what would make an example useless.

02

Specify

Requirements become a written collection or annotation guideline with worked examples and an acceptance definition, agreed before work starts.

03

Execute

Trained annotators and collectors work to that guideline, with edge cases escalated rather than guessed at.

04

Validate

Human review against the same rubric, plus consistency checks across annotators before anything is marked complete.

05

Deliver

Structured output in the schema your pipeline expects, with the guideline and review notes attached.

Quality workflow

Six gates between raw input and delivery.

The process does not change per data type. A corpus assembled across more than one category has to hold together, so the acceptance standard is shared.

01

Collection

Capture or sourcing against a defined specification, so the input is already in scope.

02

Guidelines

A written definition of correct, with worked examples and explicit edge-case handling.

03

Annotation

Trained annotators working to that guideline, escalating ambiguity rather than guessing.

04

QA

Structural and consistency checks across the batch, not just within single items.

05

Human Validation

Review by a second person against the acceptance definition.

06

Final Dataset

Structured delivery with the guideline and review notes attached.

Dataset preparation

What arrives at the end.

The output is agreed before work starts, so delivery is a handover rather than a translation exercise.

  • Labels in the schema your pipeline expects, with span offsets aligned to the source text
  • Ambiguous cases marked rather than silently resolved, so your team can see where the schema strained
  • Agreement checks across annotators reported with the delivery
  • The annotation guideline supplied with the data, including how each edge case was decided

FAQ

Text Annotated Data questions.

Classification, named entity annotation, intent labelling, sentiment annotation and span-level marking, scoped to the project.

Have a data challenge?
Let’s build the dataset behind it.

Tell us what you’re building, what data you need and the scale of your project.