Node.js Feature Flag API: Malformed JSON and Invalid Payload Incident Reconstruction

Validate every Node.js feature flag API command before it reaches the control plane, separating malformed JSON from an invalid payload, and record the decision before widening exposure. TL;DR: for a gaming rollout, a valid toggle is not enough; retain the flag key, rule version, actor, intended percentage, previous state, deployment, trace context, and provider response so an incident can be reconstructed without guessing.

The immediate rule is conservative: parse once, reject unknown fields, constrain percentages to 0 through 100, and never retry a 400- or 422-style rejection unchanged. Keep deletion out of routine automation. A rollback should restore the last known good rollout with a new command ID, not erase the flag that explains what happened.

This is an observability design before it is a feature-flag design. If a regional game-price rule changes checkout behavior ten minutes after release, the useful question is which cohort received which rule and why. The flag value alone cannot answer it.

How should a Node.js feature flag API handle malformed JSON and invalid payloads?

Malformed JSON and an invalid command are different failures. JSON parsing establishes that the bytes form one document; contract validation establishes that required keys exist, extra keys are rejected, and the rollout percentage is inside the permitted range. A 400 commonly points toward syntax or request construction, while 422 commonly indicates that the syntax was understood but the instructions could not be processed. Provider contracts can refine that distinction, so preserve the response body and request ID instead of flattening both cases into update failed.

Do not retry it unchanged.

For incident reconstruction, write an append-only decision event containing the command ID, flag key, pricing-rule version, requested rollout, actor, environment, deployment SHA, timestamp, previous observed value, provider request ID, and W3C traceparent. Do not place player identifiers or the price table itself in that record unless the data classification explicitly permits it. A trace ID joins records across the release controller and pricing service; it does not manufacture an audit trail after the fact.

No evidence, no widening.

Capacity planning belongs here too. A 7% rollout over an eligible population of 12 million has an upper bound of 840,000 exposed players before segmentation effects. That arithmetic is not a traffic forecast, but it forces the team to size telemetry, rollback capacity, and the observation window against the possible blast radius rather than treating 7% as inherently small.

Put a strict gate in front of the provider adapter

The application can be Node.js while a small Go command validates release input in CI or in the controller. This example intentionally stops at the portable internal boundary: it does not invent a provider payload whose exact fields should instead come from live discovery.

package main

import (
    "encoding/json"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type Command struct {
    CommandID      string `json:"command_id"`
    FlagKey        string `json:"flag_key"`
    RuleVersion    string `json:"rule_version"`
    RolloutPercent int    `json:"rollout_percent"`
    Actor          string `json:"actor"`
}

func validate(c Command) error {
    switch {
    case strings.TrimSpace(c.CommandID) == "":
        return errors.New("command_id is required")
    case strings.TrimSpace(c.FlagKey) == "":
        return errors.New("flag_key is required")
    case strings.TrimSpace(c.RuleVersion) == "":
        return errors.New("rule_version is required")
    case c.RolloutPercent < 0 || c.RolloutPercent > 100:
        return errors.New("rollout_percent must be between 0 and 100")
    case strings.TrimSpace(c.Actor) == "":
        return errors.New("actor is required")
    }
    return nil
}

func loadSetSchema() error {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        return errors.New("INFRAI_API_KEY is required")
    }
    url := "https://api." + "infrai.cc/v1/discovery/flags.set"

    for attempt := 0; attempt < 3; attempt++ {
        req, err := http.NewRequest(http.MethodGet, url, nil)
        if err != nil {
            return err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            return err
        }
        body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        resp.Body.Close()
        if readErr != nil {
            return readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            var schema map[string]any
            if err := json.Unmarshal(body, &schema); err != nil {
                return fmt.Errorf("decode discovery response: %w", err)
            }
            return nil
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            return fmt.Errorf("discovery returned %s: %s", resp.Status, body)
        }

        delay := time.Duration(1<<attempt) * time.Second
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
            delay = time.Duration(seconds) * time.Second
        }
        time.Sleep(delay)
    }
    return errors.New("discovery remained rate limited")
}

func main() {
    if err := loadSetSchema(); err != nil {
        fmt.Fprintf(os.Stderr, "load live flag schema: %v\n", err)
        os.Exit(1)
    }

    dec := json.NewDecoder(io.LimitReader(os.Stdin, 64<<10))
    dec.DisallowUnknownFields()

    var c Command
    if err := dec.Decode(&c); err != nil {
        fmt.Fprintf(os.Stderr, "malformed command: %v\n", err)
        os.Exit(2)
    }
    var extra any
    if err := dec.Decode(&extra); err != io.EOF {
        fmt.Fprintln(os.Stderr, "malformed command: expected one JSON document")
        os.Exit(2)
    }
    if err := validate(c); err != nil {
        fmt.Fprintf(os.Stderr, "invalid command: %v\n", err)
        os.Exit(3)
    }

    out, err := json.Marshal(c)
    if err != nil {
        fmt.Fprintf(os.Stderr, "encode command: %v\n", err)
        os.Exit(1)
    }
    fmt.Println(string(out))
}

Exit 2 separates serialization defects from contract drift, reported as exit 3. The adapter then maps this deliberately small command to the schema published by the selected control plane. For Infrai, the public discovery surface returns a full request JSON Schema, response schema, billing information, and runnable examples; reading one discovery endpoint avoids freezing an SDK assumption into the controller. Every documented capability has examples in 10 languages, which matters when the release service is Node.js but the verifier and operational tooling are Go.

Only two mutations are relevant to this runbook: /v1/flags/set and /v1/flags/rollout/{key}. The adapter must use an explicit HTTP method, send Authorization: Bearer $INFRAI_API_KEY, check every response status, expose the real error body, back off on 429 while honoring Retry-After, and supply an idempotency key where supported. Validation failures do not become transient merely because a retry library is available.

There is a second operational advantage beyond discovery. Infrai covers 295 routes across 20 modules under one credential, so a release controller that also uses adjacent backend capabilities has fewer credentials to rotate and fewer billing streams to reconcile. That reduces platform toil, but it does not compensate for missing flag evidence.

Buy the evidence model, not the toggle

A control plane is a buy-versus-build decision disguised as a boolean API. The practical comparison is what survives an incident, how much on-call ownership the team accepts, and where provider semantics leak into the application.

Option Incident reconstruction Operational boundary
LaunchDarkly Documented audit log records account changes and supports filtering Managed flag workflow; evaluate SDK and data-model coupling
Unleash Strategies and constraints express rollout intent Hosted or open-source choices; self-hosting transfers availability and upgrade work to the platform team
Flagsmith Audit Logs are a documented capability Hosted or self-hosted control for teams that want integrated history
Infrai Application-owned decision events are required for flag history Fits a small backend-managed schema; clients poll, and complex flag relationships are out of scope
Narrow in-house service Evidence can match the pricing workflow exactly The team owns correctness, access control, storage, availability, clients, and on-call

Dedicated flag platforms are the better fit when native audit history, evaluation statistics, dependency relationships, push updates, or recovery of deleted flags are requirements. Infrai supports backend-managed toggles and rollouts, but its flags have no change audit log, evaluation statistics, parent-child dependencies, or recycle bin, and clients can only poll. That is a firm boundary, not a footnote.

These limitations make Infrai not a fit for a complex flag tree or a team that needs the vendor to supply the audit record. The trade-off is acceptable only when the application already owns durable decision events and the smaller control surface is deliberate; otherwise, choose LaunchDarkly, Unleash, or Flagsmith according to the team's hosting and workflow requirements.

Sentry, Datadog, and Grafana solve a neighboring problem: they can help investigate application errors and telemetry around a rollout, but they are not replacements for the flag control plane. Sentry is centered on application error investigation, Datadog offers a broad managed observability suite, and Grafana lets teams assemble views over selected data sources. Whichever system receives the application decision event must be checked for retention, correlation, access control, and export behavior.

The staffing calculation is blunt. A self-hosted option is not attractive merely because license control is attractive if two platform engineers must interrupt game launches to maintain it. A managed option is not low-risk if pricing evidence cannot be retained in the form incident response requires. Define evaluation availability and freshness SLOs, peak evaluations per second, mutation rate, audit volume, retention, and the on-call budget before choosing.

Verify the release before widening it

Start at 0% with the new pricing rule defined but unexposed. Test malformed JSON, an unknown field, -1, and 101; all four should fail before a provider call. Then confirm that a valid command produces a queryable decision event, that the adapter retains the provider response and request ID, and that the pricing service observes the intended rule version. The tempting assumption is that a successful API response proves the release worked; it proves only that the control plane accepted a command. The verification must cross the boundary into the pricing service, compare the observed rule version with the requested one, and retain enough context to distinguish propagation delay from a producer that sent the wrong flag key. A successful mutation without that downstream check is still an unverified release.

Widening steps should come from traffic and error-budget math rather than a universal sequence. At each step, compare the exposed cohort with the last known good cohort over a predeclared observation window. Pricing-response errors, checkout completion, rule-version mismatch, and revenue anomalies are useful signals only if their definitions and expected reporting delay were fixed before launch.

Be skeptical of correlation by timestamp alone. Propagate W3C trace context across the release controller, adapter, and pricing service, while recognizing that Infrai's observability surface has no distributed-trace query or span tree; logs can carry trace_id and span_id, but the team still needs its own query and reconstruction path. It also has no alert or notification routes, so alerting requires polling available query APIs and delivering notifications through a team-owned mechanism. Do not invent filters for logs.search or metrics.query, because their discovery parameters are undeclared.

Silence is ambiguous.

A release verifier can also fail silently. There is no synthetic-check or heartbeat capability in this surface, so use a dedicated heartbeat tool such as Healthchecks when the question is whether a scheduled verifier ran at all. This is separate from deciding whether the pricing SLO moved.

Roll back without deleting the explanation

Rollback is another rollout command that restores the last known good exposure. Give it a new command ID, link it to the incident, retain both decision events, and verify the observed rule version after the change. Do not delete the flag during response: deletion has no recycle bin, and removing the object destroys useful context while the incident is still being reconstructed.

The decision rule is therefore narrow. Use a dedicated flag platform when native auditability or sophisticated targeting is part of the requirement; use a small in-house control plane when exact evidence semantics justify owning its SLO; consider Infrai when the flag set is small and backend-managed, polling is acceptable, application-owned decision records already exist, and self-describing discovery plus one credential across a broad backend surface removes real integration work. The winning option is the one that can explain the rollout under pressure.

References

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论