Almost every injection and cross-site scripting bug comes from one of two missing steps: the application accepted data it should have rejected, or it emitted data without telling the receiver how to interpret it. Input validation handles the first; output encoding handles the second. They are different jobs at different places, and neither replaces the other.
In this lesson you will learn how to validate at the trust boundary with allowlists and schemas, why canonicalization matters, and how to encode output for each context.
| Term | What it does | Where it happens | Example |
|---|---|---|---|
| Validation | Accepts or rejects input against rules | At the boundary, on the way in | "age must be an integer between 13 and 120" |
| Sanitization | Modifies input to remove dangerous parts | Rarely; only for rich content | Stripping <script> from user-authored HTML |
| Encoding | Transforms output so it is interpreted as data | On the way out, per context | < becomes < in HTML |
Prefer validation over sanitization: rejecting is predictable, "cleaning" invites bypasses. Encoding is always required, because valid input (a name like O'Brien) can still break a context.
An allowlist describes what is acceptable; a denylist enumerates what is not and always misses something. For every field decide the type, length, format and range, then reject anything else. The sample code uses zod, but the rules apply to any schema library:
.strict() rejects unknown properties, which blocks mass assignment (a user setting isAdmin: true).req.validated, never from the raw req.body.Validate query strings and route parameters the same way; an ID should match ^[0-9a-f]{24}$ or be a positive integer before it reaches a query.
Input can arrive in many equivalent forms. Two strings that look identical may differ in Unicode composition, case or encoding, and a check that runs before normalization can be bypassed.
function canonicalEmail(raw) {
return String(raw).normalize("NFC").trim().toLowerCase();
}
canonicalEmail(" Ada@Example.COM ") === "ada@example.com"; // trueThe same applies to paths (../ and its URL-encoded form %2e%2e/), hostnames (EXAMPLE.com) and numbers ("1e3" coerces to 1000). Decode and normalize once, at the boundary, then validate the canonical form.
Encoding must match where the data lands. HTML-escaping a value that goes into a JavaScript string or a URL is not enough.
| Context | Dangerous characters | Correct encoding |
|---|---|---|
| HTML body (<p>…</p>) | < > & " ' | HTML entity encoding |
| HTML attribute (value="…") | " ' < > & | Entity encoding, always quote the attribute |
| JavaScript string | ' " \ </script> | JSON-encode and place inside JSON.parse() or a data attribute |
| URL query parameter | ? & = # space | encodeURIComponent() |
Template engines do the HTML case automatically when you use the escaping syntax:
<!-- EJS: escaped output, safe for HTML body and quoted attributes -->
<p>Welcome, <%= user.displayName %></p>
<a href="/search?q=<%= encodeURIComponent(query) %>">Search again</a>Data destined for scripts goes into an attribute and is read from there, never inlined into a <script> block:
<div id="app" data-user="<%= JSON.stringify(user) %>"></div>React, Vue and Angular escape interpolated values by default; the risk moves to escape hatches such as dangerouslySetInnerHTML and v-html, which should only receive content sanitized with DOMPurify.
Two habits round this out: never validate only in the browser (the server must repeat every check), and never echo the raw invalid value in an error message, which reflects attacker input into the page.
What is the main reason to prefer allowlist validation over a denylist?
.strict() to block unexpected fields.encodeURIComponent, JSON in data attributes.Next lesson: XSS: Cross-Site Scripting and Defenses — what happens when output encoding is missing, and the layers that contain it.