Input Validation and Output Encoding

Intermediate
12 min

Input Validation and Output Encoding

Almost every injection and cross-site scripting bug comes from one of two missing steps: the application accepted data it should have rejected, or it emitted data without telling the receiver how to interpret it. Input validation handles the first; output encoding handles the second. They are different jobs at different places, and neither replaces the other.

In this lesson you will learn how to validate at the trust boundary with allowlists and schemas, why canonicalization matters, and how to encode output for each context.

Validation, Sanitization and Encoding Are Not the Same

| Term | What it does | Where it happens | Example | |---|---|---|---| | Validation | Accepts or rejects input against rules | At the boundary, on the way in | "age must be an integer between 13 and 120" | | Sanitization | Modifies input to remove dangerous parts | Rarely; only for rich content | Stripping <script> from user-authored HTML | | Encoding | Transforms output so it is interpreted as data | On the way out, per context | < becomes &lt; in HTML |

Prefer validation over sanitization: rejecting is predictable, "cleaning" invites bypasses. Encoding is always required, because valid input (a name like O'Brien) can still break a context.

Validate at the Boundary with Allowlists

An allowlist describes what is acceptable; a denylist enumerates what is not and always misses something. For every field decide the type, length, format and range, then reject anything else. The sample code uses zod, but the rules apply to any schema library:

  • .strict() rejects unknown properties, which blocks mass assignment (a user setting isAdmin: true).
  • Enumerations restrict values such as roles or statuses to a fixed list.
  • Length limits protect the database and stop ReDoS on later regex checks.
  • The handler reads from req.validated, never from the raw req.body.

Validate query strings and route parameters the same way; an ID should match ^[0-9a-f]{24}$ or be a positive integer before it reaches a query.

Canonicalize Before You Compare

Input can arrive in many equivalent forms. Two strings that look identical may differ in Unicode composition, case or encoding, and a check that runs before normalization can be bypassed.

javascript
function canonicalEmail(raw) { return String(raw).normalize("NFC").trim().toLowerCase(); } canonicalEmail(" Ada@Example.COM ") === "ada@example.com"; // true

The same applies to paths (../ and its URL-encoded form %2e%2e/), hostnames (EXAMPLE.com) and numbers ("1e3" coerces to 1000). Decode and normalize once, at the boundary, then validate the canonical form.

Encode Output for the Right Context

Encoding must match where the data lands. HTML-escaping a value that goes into a JavaScript string or a URL is not enough.

| Context | Dangerous characters | Correct encoding | |---|---|---| | HTML body (<p>…</p>) | < > & " ' | HTML entity encoding | | HTML attribute (value="…") | " ' < > & | Entity encoding, always quote the attribute | | JavaScript string | ' " \ </script> | JSON-encode and place inside JSON.parse() or a data attribute | | URL query parameter | ? & = # space | encodeURIComponent() |

Template engines do the HTML case automatically when you use the escaping syntax:

html
<!-- EJS: escaped output, safe for HTML body and quoted attributes --> <p>Welcome, <%= user.displayName %></p> <a href="/search?q=<%= encodeURIComponent(query) %>">Search again</a>

Data destined for scripts goes into an attribute and is read from there, never inlined into a <script> block:

html
<div id="app" data-user="<%= JSON.stringify(user) %>"></div>

React, Vue and Angular escape interpolated values by default; the risk moves to escape hatches such as dangerouslySetInnerHTML and v-html, which should only receive content sanitized with DOMPurify.

Two habits round this out: never validate only in the browser (the server must repeat every check), and never echo the raw invalid value in an error message, which reflects attacker input into the page.

Quick Quiz
Question 1 of 2

What is the main reason to prefer allowlist validation over a denylist?

Key Takeaways

  • Validation rejects bad input at the boundary; encoding makes good input safe on the way out; both are required.
  • Use schema-based allowlists with types, lengths, formats, enums and .strict() to block unexpected fields.
  • Canonicalize (normalize, trim, decode) once before validating or comparing.
  • Encode per context: HTML entities, encodeURIComponent, JSON in data attributes.
  • Client-side validation is a convenience, not a control; the server always re-validates.

Next lesson: XSS: Cross-Site Scripting and Defenses — what happens when output encoding is missing, and the layers that contain it.

Input Validation and Output Encoding - Cyber Security | CodeYourCraft | CodeYourCraft