Requirements
- Input consists of a series of text lines, where each well-formed line follows
<unix_timestamp> <LEVEL> <message> (for instance, 1714501203 WARN cache miss).
- Produce structured entries containing, at least,
{timestamp, level, message} for every parsed record.
- Check every line against the required layout and identify lines that are invalid.
- Once the core parser works, allow retrieval of records whose timestamps fall in the inclusive
[from_ts, to_ts] interval, as required by the OHAI variation.
- Follow-up: a log message can continue across several physical lines. Treat a line as the start of a new record only if it begins with a valid
<unix_timestamp> token; append all intervening lines to the active record's message text.
Examples
Basic input:
1714501203 INFO service started
1714501211 WARN cache miss
1714501224 ERROR request aborted
Multi-line follow-up input:
1714501230 INFO job began successfully
1714501237 WARN retrying remote call
because the first attempt timed out
1714501249 ERROR job could not finish
additional diagnostic details
are included here
For the follow-up input, the records occupy lines 1, lines 2-3, and lines 4-6, respectively.
Preparation
- Build the straightforward parser within 15 minutes, then layer in multiline handling without starting over; the design should make that enhancement natural.
- Before coding, rehearse describing a single test case aloud, since ambiguity around when a record begins should be resolved through a confirmed example.