04 β Data Type
Text Annotated Data
Structured language datasets for NLP, LLM and conversational AI.
What it is
It looks simple until two people disagree.
Text annotation looks simple until two annotators disagree about the same sentence β and then it becomes a question of definitions.
Most of the value in a language dataset comes from resolving those definitions up front: what an entity boundary includes, how ambiguity is handled, and what to do with examples that genuinely sit between two labels.
Mvaak treats the guideline as the deliverable's foundation, and measures agreement against it rather than assuming it.
Why it matters for AI
Three things that decide whether it trains well.
- 01
Disagreement is information
When two annotators split on the same example, the label set is usually underspecified. Treating that as a signal rather than an error is what stops the ambiguity reaching the model.
- 02
Boundaries are a decision, not a fact
Whether a title, a qualifier or a trailing clause belongs inside an entity span changes what the model learns. Left to individual judgement, it changes per annotator.
- 03
The awkward middle carries the weight
Clear examples teach a model little it could not infer. The examples that sit between two labels are where a dataset earns its cost β and where a guideline has to be explicit.
Applications
What this data trains.
Each of these draws on the category differently, which is why the specification is written per project rather than per service.
- NLP
- LLM Training
- Text Classification
- Entity Recognition
- Intent Classification
- Sentiment Analysis
- Conversational AI
- Language Intelligence
Collection possibilities
Where the data comes from.
Most engagements are a mix: annotate what you already hold, and collect specifically where coverage is thin.
Annotating your corpus
Most text work starts from documents, transcripts, tickets or conversations you already hold, labelled against a schema defined with your team.
Targeted collection
Where a class is rare, examples can be sourced or written specifically to cover it, so the label is not learned from a handful of instances.
Guideline development
Where the definitions are not settled, the first pass is a small annotated sample used to expose disagreement and tighten the schema before volume begins.
Annotation possibilities
Scoped to the project, not sold as a bundle.
These are the operations available for this data type. Which of them a project uses depends on what the model has to learn.
- 01Classification and tagging
- 02Named entity annotation
- 03Intent labelling
- 04Sentiment annotation
- 05Span and relation marking
- 06Human validation of edge cases
Workflow
The same five stages, every project.
Understand
We start from the model and the use case, not the file format β what the system needs to learn, and what would make an example useless.
Specify
Requirements become a written collection or annotation guideline with worked examples and an acceptance definition, agreed before work starts.
Execute
Trained annotators and collectors work to that guideline, with edge cases escalated rather than guessed at.
Validate
Human review against the same rubric, plus consistency checks across annotators before anything is marked complete.
Deliver
Structured output in the schema your pipeline expects, with the guideline and review notes attached.
Quality workflow
Six gates between raw input and delivery.
The process does not change per data type. A corpus assembled across more than one category has to hold together, so the acceptance standard is shared.
Collection
Capture or sourcing against a defined specification, so the input is already in scope.
Guidelines
A written definition of correct, with worked examples and explicit edge-case handling.
Annotation
Trained annotators working to that guideline, escalating ambiguity rather than guessing.
QA
Structural and consistency checks across the batch, not just within single items.
Human Validation
Review by a second person against the acceptance definition.
Final Dataset
Structured delivery with the guideline and review notes attached.
Dataset preparation
What arrives at the end.
The output is agreed before work starts, so delivery is a handover rather than a translation exercise.
- Labels in the schema your pipeline expects, with span offsets aligned to the source text
- Ambiguous cases marked rather than silently resolved, so your team can see where the schema strained
- Agreement checks across annotators reported with the delivery
- The annotation guideline supplied with the data, including how each edge case was decided
FAQ
Text Annotated Data questions.
Classification, named entity annotation, intent labelling, sentiment annotation and span-level marking, scoped to the project.
Other data types
- 01
Egocentric Video Data
Human-perspective video data for systems learning to understand activities, environments and interactions.
- 02
Audio Annotated Data
Structured audio datasets for speech, voice and audio intelligence.
- 03
Video Annotated Data
Annotated visual datasets for computer vision and video intelligence.
Have a data challenge?
Letβs build the dataset behind it.
Tell us what youβre building, what data you need and the scale of your project.