CSV vs JSON array vs NDJSON — picking the right record format for exports and streams

This worksheet helps engineers and designers choose between CSV, a single JSON array, and newline‑delimited JSON (NDJSON) when exchanging tabular or streamed record data by clarifying streamability, schema expression, tooling, and error‑isolation trade‑offs.


How to read this worksheet

The vertical rail to the left numbers nine compact sections. Each numbered section presents three parallel "specimens"—CSV, JSON array, NDJSON—focused on the criterion named in the section heading. Use the selector below to highlight the lane and the related row in the synthesis card when you have a single primary constraint.


Orientation & schema

CSV — row-oriented

CSV is naturally row-oriented: each line maps to a record of delimited fields. Schema is implied by header rows or documentation rather than embedded types; structural variants can be tolerated but may require consumer agreements.

JSON array — document-oriented

A single JSON array is a self-contained document of records. Schema can be expressed with JSON Schema or a shared contract, but the whole array is delivered as one atomic payload by typical transports.

NDJSON — line-delimited JSON

NDJSON writes one JSON record per newline. Each line is a complete JSON value, enabling record-by-record handling; schemas are per-record or shared out-of-band similar to CSV but with explicit typing in JSON values.


Streamability & parsing

CSV — stream-friendly with caveats

CSV is efficient to stream line by line but needs careful handling of quoted fields and line breaks within fields; incremental parsing is common in tooling but must respect quoting rules.

JSON array — whole-document parsing

A JSON array normally requires the consumer to see the full document boundary before extracting records. This simplifies framing but makes true streaming and incremental processing harder without additional framing conventions.

NDJSON — naturally incremental

NDJSON supports incremental parsing: consumers can process, persist, or forward each JSON line as it arrives. Framing is simple (newline) and works well for streaming pipelines and log-like flows.


Error isolation & recovery

CSV — partial recovery possible

When a CSV row is malformed, consumers can often skip or log the offending line and continue; however, parsing ambiguity (e.g., unmatched quotes) can block recovery unless robust parsers are used.

JSON array — atomic failure risk

A single malformed element or a syntax error typically invalidates the entire array document; recovery requires partial reads or reparsing strategies and is thus harder than line-delimited formats.

NDJSON — per‑record isolation

Because each line is an independent JSON value, a parse error affects one record only; consumers can log, skip, or quarantine the bad line and continue processing the stream.


Tooling compatibility & human editing

CSV — spreadsheet-friendly

CSV is widely supported by spreadsheets and simple editors, making it the default for human-facing exports. It is concise but loses typed structure without conventions for nested data.

JSON array — tooling for documents

JSON arrays are friendly to document-oriented tools, validators, and schema processors. Less convenient for direct spreadsheet editing, but well-suited when consumers expect structured JSON.

NDJSON — developer-friendly streams

NDJSON is handy for line-oriented tooling, command-line filters, and simple stream processors. It is editable in plain text but less friendly than CSV for spreadsheet workflows.


Cross-cutting comparisons

CSV vs NDJSON — streaming logs

For log-style, append-only streams where individual records may carry nested fields, NDJSON encodes structured values directly and isolates errors per record; CSV is lighter and spreadsheet-compatible but needs escaping conventions for nested or quoted values.

JSON array vs NDJSON — API batch responses

APIs that return a small-to-medium batch where consumers expect a single payload often prefer a JSON array for atomic delivery; for very large batches or streaming processing, NDJSON enables incremental consumption and partial progress reporting.


Synthesis — compact comparison

Trade-offs across recurring criteria (rows). Use the selector to emphasize a single primary constraint.
Criterion CSV JSON array NDJSON
Orientation Row/field grid; light and compact. Document of records; atomic payload. Line-delimited JSON records; explicit values.
Schema Implied by headers or docs; limited typing. Best for schema-driven validation tools. Schema out-of-band or per-record; typed values possible.
Streaming Streamable but quoting complicates framing. Requires buffering; not ideal for streams. Designed for incremental consumption.
Error isolation Line-level skip often possible; parser-dependent. Single failure can invalidate whole document. Per-line isolation; easy to skip or quarantine bad records.
Tooling Excellent spreadsheet support. Rich JSON tooling and validators. Good for logs, CLI tools, and streams.
Human editing Easy in spreadsheets and editors. Readable as JSON but less spreadsheet-friendly. Editable in text editors; not spreadsheet-native.
Edge cases Nested data and multiline fields are awkward. Complex structures are supported naturally. Supports nested JSON per line; newline framing matters.
Picker guidance (qualitative): small human-editable exports → CSV; large streaming feeds → NDJSON; strict-schema API batches → JSON array.

Decision aid — quick heuristics & rubric

  1. Do you need record-level recovery or incremental processing? If yes, favor NDJSON.
  2. Is maximum compatibility with spreadsheets or non-developer users the priority? If yes, favor CSV.
  3. Does the consumer expect a single atomic JSON payload or schema-validated document? If yes, favor JSON array.
  4. Will records contain nested objects or typed values you want preserved without custom escaping? Prefer JSON array or NDJSON.
  5. Is framing simple and lossless required for streaming transports? NDJSON typically simplifies framing.
Source‑evaluation rubric — a tiny self‑check before implementation:
  • Confirm consumer tooling: can they process NDJSON or do they expect CSV/JSON?
  • Decide how you will express schema (header, JSON Schema, external contract) and whether that fits the format's assumptions.
  • When many constraints conflict, prototype a small sample payload and validate with representative consumers before locking the format.
When the limitations of plain-text formats are reached (complex binary columns, enforced typed schemas, or compact columnar storage), consider binary or schema-rich alternatives conceptually — this worksheet focuses only on plain-text trade-offs.