Application security & secure coding
Security as a coding discipline rather than a product: the vulnerability classes that recur in every language, the split between authentication and authorization, cryptography fundamentals, and how the OWASP list frames risk. Interviewers ask because resisting these attacks is expected of every developer, not just of a security team.
on this pageshowhide
guide
overview
~1 minApplication security questions test whether you write code that holds up when the party on the other end is hostile. They rarely want a list of named attacks; they hand you an endpoint, a file upload, a token or a stored password and watch two things: whether you find the place where attacker-controlled bytes gain power they should not have, and whether your fix removes the whole class of bug or only the example in front of you. Resisting these attacks is expected of every developer, not only of a security team, so the subject turns up in general engineering rounds too. The subject has four sections. [Secure coding principles](/topics/found-appsec-secure-coding) is the largest: the injection family (database queries, operating-system commands, file paths, deserialized objects), [input validation](/topics/found-appsec-input-validation) and [output encoding](/topics/found-appsec-output-encoding) as the two halves of handling untrusted data, and the handling of secrets and sensitive data. [Cryptography concepts](/topics/found-appsec-cryptography) covers the primitives a developer chooses between: hashing, symmetric and asymmetric encryption, cipher modes, signatures, randomness, password storage and transport security. [Authentication versus authorization](/topics/found-appsec-authn-vs-authz) separates proving who is calling from deciding what that caller may do. The [OWASP Top 10](/topics/found-appsec-owasp-top-10) is the shared vocabulary that frames these risks, and its section asks what that list is and what it is not. Start with untrusted input: where it enters, how it reaches an interpreter, and why filtering it is weaker than keeping it apart from code. Move to access control next, because broken authorization is the failure found most often in tested applications. Take cryptography after that, at the level of which guarantee each primitive buys. Questions run from junior to principal, and the same idea often returns at several depths, so revisit a section as your level rises.
primer
A few ideas recur across nearly every section. With them in place, most questions below read as one of these applied to a new surface. - **Injection is data being read as code.** Database query injection, command injection, cross-site scripting, path traversal and unsafe deserialization look like separate topics, but each has an interpreter that lets the content of an input choose what role it plays. The durable fix changes the channel so the input can only ever be a value: bound parameters, an argument list instead of a command string, a file resolver confined to one directory, a data format with no way to name types. Stripping dangerous characters is the answer interviewers push back on: the interpreter, not you, defines what is dangerous. - **Closed-world rules beat open-world ones.** Listing what is permitted rejects anything unforeseen; listing what is forbidden lets it through. The same asymmetry settles server-side request forgery defences, caller-chosen sort columns, redirect targets and deserializable types. - **Check the value you actually use.** Many bypasses come from a gap between what was validated and what was consumed: two decoders that read the same bytes differently, a path checked as text and opened after resolution, a size check whose own arithmetic wraps around. Reduce input to one form, then validate and consume that same form. - **Identity and permission are separate decisions.** Authentication settles who is calling. Authorization is decided again for every action on every resource, from state the server owns. The typical access-control bug is a missing per-record check, not a failed login, and an unguessable identifier does not stand in for that check. - **Choose a cryptographic primitive by the property you need.** Confidentiality, integrity, origin authenticity, non-repudiation and unpredictability are distinct guarantees, and no single primitive gives all of them. - **A secret is authority handed to whoever holds it.** Its worth depends on who can read every copy and how fast it can be revoked. Rotation shortens how long a leak stays useful; it does not undo the leak. - **Every claim needs an attacker model.** Say who the attacker is, what they control and what they gain if the control fails. Blast radius, not the name of the bug, decides how urgent a finding is.
- Trust boundary
- A point where data or control passes from a party with less authority to one with more; the place validation belongs, wherever the bytes came from.
- Sink
- The place where a value is handed to an interpreter such as a query engine, a shell, a browser or a file resolver, and where injection actually happens.
- Allowlist
- A rule that enumerates what is permitted and rejects everything else, so an unforeseen input fails safe rather than slipping through.
- Canonicalization
- Reducing input to a single standard representation before checking it, so that encoded or equivalent variants cannot pass a check their decoded form would fail.
- Confused deputy
- A component with more authority than its caller that performs an action the caller chose, using its own authority instead of the caller's.
- Object-level authorization
- A permission check on the specific record being accessed, not just on the endpoint or role; its absence is the most common access-control defect.
- Bearer credential
- A token or key that grants access to whoever presents it, with no proof that the presenter is the party it was issued to.
- Authenticated encryption
- Encryption that also produces an integrity tag, so any modification of the ciphertext is detected and rejected instead of decrypted.
- Nonce
- A value used once per message under a given key; many encryption modes lose their guarantees entirely if it ever repeats.
- Work factor
- The tunable cost parameter of a password-hashing function, set so each guess is expensive for an attacker while a single login stays affordable.
- Forward secrecy
- The property that recorded traffic stays protected even if a server's long-term private key is later stolen, because session keys were ephemeral.
- Cryptographically secure generator
- A random generator designed so its future output cannot be predicted from past output; required for keys, tokens and nonces.
The four sections lean on each other, and many senior questions sit where two of them meet. **Validation and encoding are the two halves of handling untrusted data.** Validation happens where data crosses a trust boundary and is phrased in the application's own terms: this is an order quantity, this is a country code. Encoding happens where a value is handed to a parser and is phrased in that interpreter's grammar. Neither replaces the other. The injection sections, covering [operating-system commands](/topics/found-appsec-command-injection), [file paths](/topics/found-appsec-path-traversal), [deserialization](/topics/found-appsec-deserialization-risks) and database queries, are that pairing worked out for one interpreter at a time, and [integer overflow](/topics/found-appsec-integer-overflow) is the same checked-versus-used gap inside plain arithmetic. **Cryptography serves the other sections.** [Password storage](/topics/found-appsec-password-storage) is where hashing meets authentication. Signatures and message authentication are what make a self-contained session token trustworthy. [Transport security](/topics/found-appsec-tls) protects the credentials and session identifiers that access control depends on, and [secure randomness](/topics/found-appsec-secure-randomness) underlies every session identifier, reset token, key and nonce. Within cryptography, [cipher modes](/topics/found-appsec-cipher-modes-integrity) explain why encryption alone gives no integrity, and transport security composes nearly every other primitive into one handshake. **[Secrets management](/topics/found-appsec-secrets-management) ties code to operations.** A key is only as safe as every place it has been copied: source history, build output, images, logs, process environment. [Sensitive data handling](/topics/found-appsec-sensitive-data-handling) extends the same question to values held in memory and written to diagnostics. **Authorization is where many design flaws hide.** The confused deputy appears both in server-side request forgery and in services that call each other with their own authority rather than the caller's. No input filter catches either; they are fixed by deciding whose authority a request carries. **The top-ten list sits on top as a map.** Its categories roll up the classes taught in the other sections. Use it to check a design's coverage, not as evidence that the design is safe.
- Input Validation & Sanitization →
Where untrusted data enters and what a check at that boundary can and cannot promise; every injection topic builds on it.
- Output Encoding & Injection Theory →
The other half of handling untrusted data: why safety depends on the destination grammar, which reframes every injection class that follows.
- Authentication vs Authorization →
Broken access control is the most common real-world failure; learn to separate identity from permission before any token or session detail.
- Cryptographic Hashing →
The simplest primitive and the most misused one; its properties are needed for password storage, signatures and integrity checks.
- Symmetric Encryption →
Establishes what encryption does and does not guarantee, the ground for cipher modes, key distribution and transport security.
- OWASP Top 10 →
Read last, once the classes are familiar, to see how the industry groups them and how to use the list on a design.
Answering an injection question with character filtering or escaping when the interpreter offers a channel that keeps data and structure apart.
Treating a logged-in user as an authorized one: checking the role on the endpoint but not whether this caller may touch the record requested.
Storing passwords with a fast general-purpose hash, salted or not; the problem is how many guesses per second an attacker gets, not reversibility.
Encrypting without integrity protection, or reusing a nonce under the same key; either can break a scheme whose cipher is perfectly sound.
Calling hashed personal identifiers anonymised; a deterministic digest is a stable pseudonym that still links records and can be enumerated.
Deleting a leaked credential from the repository and calling it fixed, instead of rotating it while every existing copy is still valid.
Presenting the OWASP Top 10 as a test plan, as though covering its categories proved an application secure.
Generating tokens or keys with the ordinary random generator because its output looks random; unpredictability, not distribution, is the requirement.
The same few choices come back across the hub, and naming the one you are making usually earns more than the fix itself. - **Structural fix versus filter.** Changing the channel so input cannot become syntax removes a class of bug but may cost features, such as shell pipelines or caller-chosen query structure. A filter keeps the feature and leaves an open-ended set of bypasses. Say which you chose and what you gave up. - **Strict validation versus legitimate variety.** Allowlists fail safe but reject real inputs nobody foresaw: names, addresses, international text. Free-form fields shift the burden onto correct encoding at every place the value is used. - **Self-contained tokens versus server-side references.** A token the service can verify alone saves a lookup but freezes a decision at issue time; a reference costs a lookup and can be revoked at once. Unless the token is bound to a key, neither limits what a thief can do with a stolen copy. - **Password cost versus login capacity.** Every unit of work an attacker pays per guess, your servers pay per login. Tune to the highest cost your capacity allows, and throttle attempts, since the costly check is itself a denial-of-service lever. - **Expressive versus auditable access models.** Roles are easy to review; attribute and relationship models express tenancy, sharing and delegation but turn the audit question, which users can read this record, into real work. - **Compatibility versus attack surface.** Every legacy protocol version or cipher suite left enabled is one an active attacker can try to steer a connection toward.
Several shapes appear under different names across the sections; recognising them is how you place an unfamiliar question quickly. - **Separate the channel.** Parameter binding, argument lists, confined file resolvers and type-free data formats all give data a path on which it cannot be parsed as instructions. - **Enumerate instead of filter.** Translate a caller's choice into a value the code owns: sort columns, file names, redirect targets, outbound destinations, deserializable types. Anything outside the table is rejected. - **Recompute permission from state the server owns.** The request names an object; the server decides, every time, whether this caller may act on it. - **Carry the caller's authority, not your own.** A downstream call presents a scoped credential for the original principal, which closes confused-deputy paths by construction. - **Bind integrity to what you rely on.** Authenticated encryption, signatures over the exact bytes you will interpret, and a check over the whole negotiation of a protocol all stop an attacker editing what you trust. - **Plan for revocation from the start.** Versioned keys, two valid generations during a change, and short credential lifetimes turn a leak into a routine rotation instead of an outage.
explore
- Secure Coding Principles45 questions
- Input Validation & Sanitization5 questions
- Output Encoding & Injection Theory4 questions
- SQL Injection5 questions
- OS Command Injection5 questions
- Path Traversal6 questions
- Insecure Deserialization Risks5 questions
- Secrets Management6 questions
- Sensitive Data in Memory & Transit5 questions
- Integer Overflow as a Security Bug4 questions
- OWASP Top 105 questions
- Authentication vs Authorization5 questions
- Cryptography Concepts37 questions
- Cryptographic Hashing5 questions
- Symmetric Encryption4 questions
- Asymmetric Encryption5 questions
- Cipher Modes & Authenticated Encryption5 questions
- Digital Signatures4 questions
- Secure Randomness5 questions
- Password Storage4 questions
- Transport Layer Security5 questions
- AI Engineerroleanchors this topic
- AI Red Teamingroleanchors this topic
- API Designskillanchors this topic
- Android Developerroleanchors this topic
- Backend Developerroleanchors this topic
- Blockchain Developerroleanchors this topic
- Cyber Security Expertroleanchors this topic
- Data Engineerroleanchors this topic
- DevOps / SRE Engineerroleanchors this topic
- DevSecOps Engineerroleanchors this topic
- Frontend Developerroleanchors this topic
- Full Stack Developerroleanchors this topic
- GraphQLskillanchors this topic
- Java Backend Developerroleanchors this topic
- Java SDETroleanchors this topic
- Kotlin Backend Developerroleanchors this topic
- Machine Learning Engineerroleanchors this topic
- QA Engineerroleanchors this topic
- Software Architectroleanchors this topic
- iOS Developerroleanchors this topic
- Computer Scienceskill
- SQLskill
questions
92 · 4 sectionsWhat is the practical difference between launching an external program by handing the operating system a single command-line string versus an explicit list of arguments, and why does the second remove a whole bug class rather than reducing it?
basics
~20 sA command string is parsed by a shell, so user bytes can become syntax. An argument list is handed to the process-creation call as an array — no parsing step exists, so a value containing ; rm -rf / is just an odd argument. You lose pipes, redirection and globbing, and must do them yourself.
Explain the difference between allowlist (positive) and denylist (negative) input validation, why one of them is structurally stronger rather than merely better practice, and where the stronger one stops working.
basics
~20 sA denylist enumerates the attacker's moves — an open-ended, growing set — and fails open on anything unforeseen. An allowlist enumerates the application's own domain — finite and known — and fails closed. The difference is open-world versus closed-world, not optimism versus pessimism.
A size check is written as `if (offset + length > buffer_size) reject;` where both `offset` and `length` come from an untrusted request. Explain why this check is unsafe on fixed-width integers and how you would rewrite it so it is correct for the whole input range.
basics
~20 sThe check itself does the overflowing arithmetic: if offset + length wraps, the sum becomes small and the guard passes for values that are far out of range. Rewrite it so nothing can wrap — compare by subtraction against the known limit, or use a checked-addition operation that fails on overflow.
The standard advice for keeping credentials out of a codebase is to inject them as process environment variables at deployment time. What does that injection actually guarantee, what does it not guarantee, and what does it imply about rotating the value?
basics
~20 sIt separates the credential from the artifact, so one build runs in many environments and the value is not in version control. It does not encrypt anything, does not narrow who on the host can read it, and it is read once at start — so rotation needs a restart.
Of all the places a secret can end up by accident, diagnostic output — application logs, error messages, crash reports and telemetry — is the one that causes the most incidents. Explain why that surface is worse than a copy sitting in process memory, list the mechanisms by which secrets reach it without anyone writing a log statement containing one, and describe the strongest structural fix.
basics
~20 sA memory copy lives milliseconds behind a process boundary; a log line lives for years, replicated, indexed, shipped to third parties and readable by many people. Secrets arrive there through automatic serialisation, exception messages and URLs — not through deliberate logging.
The OWASP Top 10 is often treated as a checklist to test an application against. Explain what the list actually is, how its entries are derived and ranked, and what a team is getting wrong when it says "we tested for the Top 10, so we are secure".
basics
~20 sAn awareness document: ten broad risk categories, each rolling up many underlying weakness types, ranked mainly by how often they appear in tested applications weighted by exploitability and impact. It is shared vocabulary and a coverage prompt, not a verification standard or a sufficient requirements list.
Server-side request forgery earned its own entry in the OWASP Top 10. Explain the trust property it violates, and why blocking private and loopback address ranges is a weaker defence than an allow-list of permitted destinations.
basics
~20 sThe server becomes a confused deputy: it makes a request an attacker chose, but with the server's network position and credentials. Blocking bad addresses is an open-world deny-list — redirects, DNS re-resolution and address encodings keep producing new bypasses — while an enumerated allow-list of destinations is closed and finite.
The OWASP Top 10 added "Insecure Design" as a category distinct from "Security Misconfiguration" and from implementation bugs. What distinguishes a design flaw from an implementation flaw, and why does the distinction change how a team responds?
basics
~20 sAn implementation flaw is a control that exists but is written wrong; a design flaw is a control that was never specified, so there is no correct code to write. Design flaws survive perfect code, perfect libraries and clean scans — they are fixed by changing requirements or the flow, not by patching a line.
Automated scanners reliably report injection findings and outdated-dependency findings, but they mostly miss the case where an authenticated request returns a record belonging to a different customer — which is exactly the shape the OWASP Top 10 ranks first by incidence. Explain the testing-theory property that separates those categories, and what it implies about how the Top 10's own numbers should be read.
basics
~20 sTesting needs an oracle: a way to judge wrongness from the observation alone. Injection and vulnerable components have one. An authenticated 200 carrying someone else's record does not — ownership is a domain fact absent from the artefact under test.
You inherit a large application portfolio and a backlog of security findings spread across the OWASP Top 10 categories. How do you decide the order of work, and why is the list's own ranking a poor prioritisation input?
basics
~20 sRank by your own risk: reachability of the sink, privilege of the executing principal, sensitivity of the data reached, and the cost of the durable fix. The list's ranking is an industry incidence statistic collected from other people's applications, so it says nothing about which of your findings an attacker can actually reach.
Authentication and authorization are routinely bundled together as 'auth'. Define each precisely, explain why they are kept as separate concerns in a system's design, and name the failure that appears when they are conflated.
basics
~20 sAuthentication establishes who the principal is, from evidence, with some level of confidence. Authorization decides whether that principal may perform this action on this resource, in this context. Conflating them produces 'logged in, therefore allowed' - the most common access-control bug.
An endpoint returns an invoice given its identifier, and any logged-in user who supplies another customer's identifier receives their invoice. Explain why this class of defect is so common, state the enforcement rule that prevents it, and say why switching to random unguessable identifiers is not the fix.
basics
~20 sFrameworks give you endpoint-level checks for free, but per-record checks need domain knowledge, so they get omitted. The rule: every request re-derives permission from server-owned state - the request says which object, never whether it is allowed. Unguessable identifiers are obscurity; identifiers leak, and then there is no check at all.
Explain the confused-deputy problem in access control, give two examples from different layers of a system, and say what structurally prevents it rather than what patches each instance.
basics
~20 sA component with more authority than its caller performs an action the caller chose, using its own ambient authority instead of the caller's. The structural fix is to carry authority with the request - a scoped, audience-bound credential representing the original principal - rather than attaching it to the deputy's identity.
A credential presented to a service is either a reference to state the issuer holds, or a self-contained assertion the verifier evaluates on its own. Where does the authority live in each shape, what exactly does an assertion freeze at the moment it is minted, and why does the choice between the two shapes not change what happens when the credential is stolen?
basics
~20 sA credential is either a reference to the issuer's state or a self-contained assertion the verifier evaluates alone. An assertion is a frozen authorization decision from issuance time. Both are bearer credentials — possession is authority — unless bound to a key.
Compare role-based, attribute-based and relationship-based access control as ways to express permissions, and describe how you would choose between them for a multi-tenant product that supports sharing and delegation.
basics
~20 sRoles are a closed, auditable set but explode when decisions depend on data. Attribute policies are expressive but make 'who can see this' hard to answer. Relationship models express sharing and inheritance naturally at the cost of consistency. Real systems combine: roles for coarse operations, relationships for object-level, attributes as conditions.
Public-key encryption is usually taught as 'encrypt with the public key, decrypt with the private key'. Give the definition that explains what asymmetric encryption actually solves, and why calling it 'a stronger shared password' is the wrong frame.
basics
~20 sA linked key pair: the public key only encrypts and may be published, the private key alone decrypts. It removes the need for a pre-shared secret, so it solves key distribution, not strength. A ciphertext still proves nothing about who produced it.
Why is Electronic Codebook (ECB) mode considered unusable for general-purpose encryption, and what property must any acceptable mode have that ECB lacks?
basics
~20 sECB encrypts each block independently with no randomization, so identical plaintext blocks give identical ciphertext blocks — it leaks structure, repetition and equality, and blocks can be reordered or spliced. Any usable mode must be randomized: same plaintext, different ciphertext each time.
A candidate defines a digital signature as "encrypting the hash with your private key". Give a definition that also holds for signature schemes involving no encryption at all, and state precisely which security property a signature provides that a shared-key message authentication code (MAC) cannot.
basics
~20 sSigning runs the message and a private key through a signing algorithm to produce a value anyone can check with the matching public key. It gives integrity and origin authenticity. A shared-key MAC gives both too, but not non-repudiation: either key holder could have produced the tag.
A team exports a customer table to an analytics partner and replaces the email column with its SHA-256 digest, reporting that the export is now "anonymised" and safe to share. Explain what a digest of a personal identifier actually gives you, how hashing differs from encryption and from encoding such as Base64, and what you would do instead.
basics
~20 sHashing is unkeyed and deterministic, so a hashed identifier is a stable pseudonym, not anonymous data: it still singles out people and joins across datasets, and small identifier spaces enumerate. Encryption is keyed and reversible; encoding is neither and protects nothing.
A team stores user passwords as SHA-256 digests and argues that SHA-256 is cryptographically strong and irreversible. Explain why that reasoning is wrong for passwords, and what property a password-storage function must have instead.
basics
~20 sPasswords are low-entropy, so nobody inverts the hash — they guess candidates and hash them. A general-purpose hash is designed to be fast, so an attacker tests billions of guesses per second. A password function must be deliberately slow and hardware-hostile, with a per-user salt and a tunable cost.