Zero is not missing data: modelling "measured", "not provided" and "unreadable"
In a scoring pipeline, an exception from a parser is a visible failure: the job fails, someone gets an alert. A parser that returns 0 when it could not read the input is an invisible one. The value has the right type, passes validation, and downstream rules treat it as a fact. In credit scoring that fact is often favourable: zero entries in a debtor registry, zero arrears, zero inquiries.
Example: an external engine returns a summary saying the applicant has no entries in a debtor registry, while the raw registry response attached to the same call contains an active entry. The summary is wrong because the parser did not match the response structure and fell back to defaults. Nothing in the logs indicates a problem.
How the zero gets in
Mostly through code that is reasonable line by line. PHP makes it easy to replace "no value" with zero:
$debts = $response['summary']['count'] ?? 0;
$amount = (int) $row['amount']; // null → 0, "" → 0, "abc" → 0
$total = array_sum($items); // [] → 0
And through constructors with defaults:
final class RegistryResult
{
public function __construct(
public int $negativeCount = 0,
public float $negativeAmount = 0.0,
) {}
}
Combine them and a parser that misses the structure of a response (a new schema version, a different XML namespace, a renamed key) does not fail. It finds nothing, and "nothing" becomes the default value. Tests built on empty or minimal responses pass, because for this parser every response looks empty.
The habit comes from treating null as a crash risk. Tony Hoare called the null reference his "billion-dollar mistake", and a lot of defensive code since then aims to make nulls disappear. The result is often worse: a loud failure replaced with a quiet wrong answer.
Three meanings of "nothing"
For data describing a person or a company, an empty result can mean three different things:
- Measured, result is zero. The registry answered and there are no entries. This is information.
- Not provided. The source was never requested or never delivered. This is a gap.
- Unreadable. The source arrived, but the parser could not interpret it. This is a failure.
Only the first may become 0. The other two must reach the decision layer as distinct states, and a rule that depends on them should return "cannot assess" instead of a score.
enum Presence
{
case Measured;
case NotProvided;
case Unreadable;
}
final readonly class Feature
{
private function __construct(
public Presence $presence,
public ?float $value,
) {}
public static function measured(float $value): self
{
return new self(Presence::Measured, $value);
}
public static function notProvided(): self
{
return new self(Presence::NotProvided, null);
}
public static function unreadable(): self
{
return new self(Presence::Unreadable, null);
}
}
The parser decides which state applies and does not guess:
final class RegistryParser
{
public function negativeCount(?array $response): Feature
{
if ($response === null) {
return Feature::notProvided();
}
$count = $response['summary']['count'] ?? null;
if (!is_int($count) || $count < 0) {
return Feature::unreadable();
}
return Feature::measured($count);
}
}
The rule handles every state explicitly. With match over an enum, a forgotten case raises UnhandledMatchError at runtime instead of falling through to a default:
enum Decision
{
case Pass;
case Reject;
case CannotAssess;
}
function debtorRegistryRule(Feature $negatives): Decision
{
return match ($negatives->presence) {
Presence::Measured => $negatives->value > 0 ? Decision::Reject : Decision::Pass,
Presence::NotProvided, Presence::Unreadable => Decision::CannotAssess,
};
}
The same distinction belongs in storage. A nullable column alone does not separate "not provided" from "unreadable"; store the state next to the value (for example negative_count integer null plus negative_count_status text not null), so reports and later analyses can filter on it.
Trade-offs: every consumer now has to handle three states, and the product needs a defined path for "cannot assess" (manual review, a request for documents, a retry). That is extra work, and it is the point: the decision about missing data moves from an accidental default to an explicit business rule. For fields where a missing value really is equivalent to zero, such as an optional discount, this machinery is unnecessary.
Tests that catch a parser returning zeros
A test with an empty response proves nothing, because a broken parser passes it. The useful fixture is a real, anonymised response that contains a positive result, with an assertion that the result is found:
public function test_detects_entry_in_registry_response(): void
{
$json = file_get_contents(__DIR__ . '/fixtures/registry_with_active_entry.json');
$response = json_decode($json, true, flags: JSON_THROW_ON_ERROR);
$feature = (new RegistryParser())->negativeCount($response);
$this->assertSame(Presence::Measured, $feature->presence);
$this->assertGreaterThan(0, $feature->value);
}
public function test_unknown_structure_is_unreadable_not_zero(): void
{
$feature = (new RegistryParser())->negativeCount(['v2' => ['items' => []]]);
$this->assertSame(Presence::Unreadable, $feature->presence);
}
Add a new fixture whenever the provider changes its schema. One non-empty sample per format catches more than many happy-path tests on synthetic data.
Checking fields across many cases
The second check needs no knowledge of the parser. Compare the same field across several reports about different subjects. A field that has the same value in every report is probably padding or a broken mapping, not a measurement. Typical candidates: a risk flag that is always set, a "has a registered business" marker that is always false, a category that appears in every report regardless of input.
/**
* @param list<array<string, mixed>> $reports
* @return list<string> fields with a single value across all reports
*/
function constantFields(array $reports): array
{
if (count($reports) < 2) {
return [];
}
$fields = array_keys($reports[0]);
return array_values(array_filter(
$fields,
static fn (string $field): bool => count(array_unique(array_map(
static fn (array $report): string => json_encode($report[$field] ?? null),
$reports,
))) === 1,
));
}
Run it on five to ten real cases before trusting a new data source. The distribution of a field's values tells you more about its quality than its name.
The same applies to third-party extractions in general. Structured JSON with readable field names is someone's interpretation of the source, including their defaults and mapping errors. Before building conclusions on it (for example, "all obligations are non-bank"), compare a few cases with the raw source document.
Checklist
- Search for
?? 0,(int)casts and numeric constructor defaults in parsers and DTOs. - Model "measured", "not provided" and "unreadable" as separate states, in code and in storage.
- Make rules return "cannot assess" for the last two, and define what happens next.
- Keep at least one non-empty, anonymised fixture per source format.
- Check value distributions across real cases before relying on a field.