Extrakce popisných dat z narativů

Abstract

The thesis focuses on transforming selected parts of unstructured route narratives into a structured representation. The implemented system is a process composed of text preprocessing, descriptivedata detection, and subsequent synthetic sentence generation. Preprocessing uses a language model together with a separate preprocessing tool authored by Marek Žáček [1]. The core contribution is a rule-based design for detecting objects, colors, sizes, verbs, and object sequences from dependency structures. The main output is a structured Prolog-fact representation suitable for deterministic machine processing and downstream corpus generation. The thesis also provides formal algorithm descriptions and consistency verification via a round-trip test. Beyond the assignment scope, the repository includes an additional experimental extension.

Description

Delayed publication

Výsledky práce jsou v připravovaném časopiseckém článku a v příspěvku na konferenci.

Available after

2028-12-31

Subject(s)

natural language processing, structured data representation, information extraction, Prolog, syntactic analysis, dependency parsing

Citation