View all services
Talk to QA Advisor
Browse the Knowledge Hub101 resources
/Interview Questions/Performance Testing

Question Bank

82 performance testing interview questions with answers

Tool agnostic by design, so this covers the discipline rather than one product. Eighty-two questions across test types, non functional requirements, workload modelling, scripting concerns including correlation and pacing, environments and data volume, execution and ramp strategy, results analysis and the statistics that get misused, bottleneck diagnosis, monitoring and APM, scalability and reporting. Graded from fresher to lead, with the model answer, the follow-up to expect, and the trap that costs candidates the round.

82questions/4experience levels/13topics/Freedownload, no sign-up

All 82 questions, with model answers

Filter by level or topic, search the full text, and download the whole bank to revise offline.

Last updated

Experience level

Topic

Showing 82 of 82 questions

Q1FresherFundamentals

What is performance testing and why is it done?

What they are assessing

Basic orientation.

Model answer

Performance testing measures how a system behaves under a given workload, in terms of responsiveness, throughput, resource usage and stability. It is done to answer questions the functional tests cannot: will the system handle the expected load, where does it break, does it stay stable over time, and what will it cost to run. The purpose is not to produce a pass or fail on a single number but to characterise behaviour so capacity and risk decisions can be made with evidence.

Q2FresherFundamentals

What is the difference between response time, latency and throughput?

What they are assessing

Core vocabulary, often used loosely.

Model answer

Response time is the total time from request to complete response as the client experiences it. Latency is usually the network delay component, the time before the response begins to arrive, though the term is frequently used loosely to mean response time. Throughput is the rate of work completed, measured in requests or transactions per second. The relationship matters: throughput can stay flat while response time climbs, which is the signature of a saturated system, and reporting only throughput hides that completely.

Likely follow-up

If throughput plateaus but load keeps increasing, what is happening?

Q3FresherFundamentals

What is the difference between concurrent users and virtual users?

What they are assessing

A distinction that drives bad test design when misunderstood.

Model answer

Virtual users are the simulated threads your tool runs. Concurrent users usually means users with an active session, and simultaneous users means users issuing a request at the same instant, which is a much smaller number. A thousand people logged in might generate only thirty simultaneous requests, because real users spend most of their time reading rather than clicking. Conflating these is how teams end up load testing at ten times the real demand and rejecting a system that would have been fine.

Trap to avoid

Treating registered users or daily active users as concurrent users. The conversion from business numbers to concurrency is where most workload models go wrong.

Q4Mid-levelFundamentals

What is the difference between performance testing, performance engineering and tuning?

What they are assessing

Whether the candidate sees the wider discipline.

Model answer

Performance testing measures and reports. Performance engineering is the broader practice of building performance into the system from design onwards, including architecture review, capacity planning, code-level profiling and setting requirements before anything is built. Tuning is the act of changing configuration or code to improve the result. A performance tester who only runs tests and hands over a report has limited value, because the hard part is diagnosing the cause and influencing the fix.

Q5Mid-levelFundamentals

When in the lifecycle should performance testing start?

What they are assessing

Shift-left thinking with a realistic view.

Model answer

Requirements should be defined at the same time as functional requirements, because a target set after the system is built is usually just a description of what it already does. Component-level testing of critical services can begin as soon as they exist, which catches algorithmic and query problems while they are cheap. Full end-to-end testing needs a stable, representative environment, so it arrives later by necessity. The failure pattern worth naming is performance testing scheduled for the last two weeks before release, where any finding is either ignored or delays the launch.

Likely follow-up

What can you usefully test before the full system exists?

Q6FresherTest types

What is load testing?

What they are assessing

The most basic type.

Model answer

Running the system at the expected workload, typically peak expected load, to verify it meets its response time and throughput requirements and remains stable. It answers whether the system handles what we expect it to handle, which is the baseline question before any other type is worth running.

Q7FresherTest types

What is the difference between load, stress and spike testing?

What they are assessing

The distinction interviewers ask for most often.

Model answer

Load testing runs at expected demand to verify requirements are met. Stress testing deliberately exceeds it, increasing load until the system degrades or fails, to find the breaking point and observe how it fails. Spike testing applies a sudden sharp increase and then removes it, to see whether the system copes with an abrupt surge and, just as importantly, whether it recovers afterwards. Recovery is the part people forget: a system that survives the spike but never returns to normal response times has still failed.

Likely follow-up

What should you observe during the recovery period after a spike?

Q8Mid-levelTest types

What is soak testing and what does it find that load testing does not?

What they are assessing

Understanding of time-dependent failures.

Model answer

Soak or endurance testing runs a moderate, realistic load for an extended period, typically eight to seventy-two hours. It finds problems that only appear with time rather than with volume: memory leaks, connection pool exhaustion, file handle leaks, log files filling a disk, database connection growth, and session or cache stores that grow without bound. A one-hour load test will pass cleanly on a system that falls over after thirty hours, which is exactly the failure that then happens in production on a long weekend.

Trap to avoid

Describing soak testing as a longer load test at higher volume. The point is duration at realistic load, not more load.

Q9Mid-levelTest types

What is volume testing?

What they are assessing

A type often confused with load testing.

Model answer

Testing with a large amount of data rather than a large number of users, to see how the system behaves when the database holds production-scale volumes. It finds missing indexes, queries that scan rather than seek, reports that time out, and pagination that degrades. It is distinct from load testing because a system can be perfectly fast with ten concurrent users and a table of fifty million rows, or fast with an empty database and a thousand users, and fail only when both are true.

Q10Mid-levelTest types

What is scalability testing and what are the two directions?

What they are assessing

Capacity thinking.

Model answer

Testing how the system responds to added resources, to establish whether it scales and how efficiently. Vertical scaling adds resources to one machine, horizontal scaling adds more machines. The useful output is the scaling factor: doubling the application servers should ideally double throughput, and the gap between that and reality tells you where the shared bottleneck is, usually the database or a lock. A system that gains nothing from a second server has a constraint that no amount of hardware will fix.

Likely follow-up

What would you conclude if adding servers increases throughput by only ten percent?

Q11SeniorTest types

What is a baseline and why does it matter?

What they are assessing

Comparative method.

Model answer

A recorded set of results under defined conditions that later runs are compared against. It matters because almost every performance question is comparative rather than absolute: is this release slower than the last, did the fix help, does the new infrastructure perform better. Without a baseline, a response time of 1.2 seconds is a number nobody can interpret. The discipline required is that the baseline must be re-measured whenever the environment or workload changes, otherwise you are comparing against conditions that no longer exist.

Q12Mid-levelRequirements

What makes a good non functional requirement for performance?

What they are assessing

Whether the candidate can turn a vague ask into something testable.

Model answer

It states the transaction, the load condition, the metric with a percentile, and the target. For example: the search results page returns within 2 seconds at the 95th percentile with 500 concurrent users during a one-hour peak. Each element matters. Without the load condition the number is meaningless. Without the percentile it is ambiguous. Without naming the transaction it cannot be measured. A requirement reading the system should be fast is not a requirement, and turning it into one is usually the first real task on a performance engagement.

Likely follow-up

How do you get those numbers when the business cannot provide them?

Q13SeniorRequirements

The business cannot tell you the expected load. How do you proceed?

What they are assessing

Practical problem solving, since this is the normal situation.

Model answer

I derive it rather than wait. For an existing system, production access logs or APM data give actual traffic patterns, peak hours and transaction mix, which is the best source available. For a new system, I work from business projections: expected customers, transactions per customer per day, the shape of the daily curve, and known seasonal peaks, then convert to concurrency with an assumed think time. Where a figure is genuinely unknown I state it as an assumption in writing and test a range rather than a point, so the result says the system supports up to X rather than claiming to validate a number nobody agreed.

Trap to avoid

Refusing to start until requirements are given. The requirement will not arrive, and the engineer who derives a defensible model and documents the assumptions is the one who is useful.

Q14SeniorRequirements

Why use percentiles rather than averages?

What they are assessing

Statistical literacy, which is where many candidates fall down.

Model answer

Response time distributions are skewed with a long tail, so the mean is dragged by outliers and simultaneously hides them. An average of 1.2 seconds is consistent with every user getting 1.2 seconds, or with ninety percent getting 0.4 and ten percent getting 9. Those are very different products. Percentiles describe the distribution: the 95th tells you what the slowest one in twenty experiences, and the 99th matters on high-volume systems because it still represents a large number of real people. I report median, 90th, 95th and 99th together, because each answers a different question.

Likely follow-up

On a system serving ten million requests a day, why does the 99th percentile matter?

Q15SeniorRequirements

What is an SLA, an SLO and an SLI?

What they are assessing

Vocabulary from the operations side, increasingly expected.

Model answer

An SLI is the indicator, the thing actually measured, such as the 95th percentile latency of the checkout endpoint. An SLO is the internal objective for that indicator, such as under 500 milliseconds for 99.9 percent of a rolling month. An SLA is the contractual commitment to a customer, usually looser than the SLO so there is margin before a breach has commercial consequences. For a performance engineer the relevance is that these give targets grounded in production behaviour rather than invented for a test, and testing against the SLO rather than the SLA is the right choice.

Q16Mid-levelWorkload modelling

What is a workload model and what goes into one?

What they are assessing

The single most important design artefact in performance testing.

Model answer

A description of the load to be simulated: which business transactions, in what proportions, at what rate, with what think time, by what user types, over what duration and ramp profile. It is derived from production data where possible. Everything else in the test depends on it, because a technically flawless test running the wrong mix of transactions produces confident and useless results. If I could keep only one artefact from a performance engagement, it would be the workload model and its justification.

Likely follow-up

Where do you get the transaction mix for an existing system?

Q17Mid-levelWorkload modelling

What is think time and why does it matter?

What they are assessing

A parameter that changes results by an order of magnitude.

Model answer

The pause between a user receiving a response and issuing their next request, simulating reading and deciding. It matters because removing it turns a realistic simulation into a throughput maximisation exercise: a hundred users with no think time can generate more load than ten thousand real ones. Without think time you are not testing the stated user count, you are testing some much larger notional population, and the concurrency figure in your report is wrong. It should also be randomised around a mean rather than fixed, since identical pauses create artificial synchronised waves of requests.

Trap to avoid

Removing think time to reach a throughput target. It produces a test that hits the number while modelling nothing real, and the report then claims a user count the test never represented.

Q18SeniorWorkload modelling

What is pacing and how does it differ from think time?

What they are assessing

A distinction that separates practitioners.

Model answer

Think time is the pause within an iteration, between steps. Pacing controls the rate at which whole iterations start, so each virtual user begins a new transaction at a fixed interval regardless of how long the previous one took. Pacing is what gives you a controlled, predictable throughput, which is why it is the right mechanism when the requirement is expressed as transactions per hour rather than as a user count. Without pacing, throughput becomes a function of response time, so as the system slows your applied load drops and the test quietly stops applying the pressure you intended.

Likely follow-up

Why does throughput falling as the system slows make a stress test misleading?

Q19SeniorWorkload modelling

How do you decide which transactions to include in a test?

What they are assessing

Prioritisation.

Model answer

By combining frequency and risk rather than either alone. Production data gives the highest-volume transactions, which must be included because they generate the load. Then I add the business-critical ones even if they are infrequent, because a slow quarterly report matters to the people who run it, and the resource-heavy ones, since a low-volume transaction doing a full table scan can dominate database load. The aim is a mix that reproduces the system-level resource profile, not a complete inventory of features. Twenty transactions covering eighty percent of load is a better test than two hundred covering everything.

Q20Mid-levelScripting

What is correlation and why is it necessary?

What they are assessing

The most important scripting concept.

Model answer

Capturing a dynamic value from a server response and using it in a subsequent request. It is necessary because recorded scripts contain values that were valid only for the recorded session: session identifiers, CSRF tokens, view state, order numbers and similar. Replayed verbatim they are rejected or, worse, silently accepted while the system does nothing useful. Correlation is the difference between a script that produces load and one that produces a stream of errors, and it is the single largest part of scripting effort on most applications.

Likely follow-up

How do you identify which values need correlating?

Q21Mid-levelScripting

What is parameterisation and why does it matter?

What they are assessing

Data handling in scripts.

Model answer

Replacing hardcoded values in a script with data drawn from a file or generator, so each virtual user operates on different data. It matters for realism and for correctness of the result. Every user searching the same term means the database serves everything from cache after the first request, and the test reports response times that will never occur in production. The same applies to logging in as one account, which serialises on row locks and measures contention rather than performance.

Trap to avoid

Running the whole test with one user account and one data set. It is the most common reason a test produces results far better than production.

Q22SeniorScripting

How do you decide between protocol-level and browser-level performance testing?

What they are assessing

A current and genuinely debated question.

Model answer

Protocol-level scripting, which is what JMeter, Gatling and k6 do by default, sends HTTP requests directly. It is enormously more efficient, so thousands of virtual users run on modest hardware, and it is the right tool for loading the server. What it does not measure is client-side cost: JavaScript execution, rendering, and the sequencing a real browser performs. Browser-level tools drive actual browsers and measure what the user experiences, but each virtual user needs a browser, so the resource cost is roughly two orders of magnitude higher. The usual answer is both: protocol-level to generate load, and a small number of real browser sessions running concurrently to measure experienced page load under that load.

Likely follow-up

Why can a server-side response time of 200 milliseconds still give a 6 second page load?

Q23SeniorScripting

How do you handle authentication in a performance script?

What they are assessing

A practical problem with several wrong answers.

Model answer

Through the real authentication flow, with correlated tokens and a pool of distinct accounts, because authentication itself is often a bottleneck and skipping it removes real load. Where tokens are short-lived, the script must handle refresh rather than failing partway through a soak test, which is a common cause of a run that degrades after an hour for reasons unrelated to the system. For protocols using a third-party identity provider, I would usually obtain tokens outside the test and feed them in, both because the provider is not ours to load test and because doing so without permission can breach their terms.

Q24SeniorScripting

What makes a performance script maintainable?

What they are assessing

Engineering discipline, often absent in performance work.

Model answer

The same things as any code, which performance scripts are frequently not treated as. Version control rather than files on an engineer laptop. Parameterised environment configuration so the same script runs anywhere. Modular transactions rather than one enormous recorded flow. Correlation handled in reusable functions rather than copied regular expressions. Meaningful transaction names, since those become the report. And assertions on response content, not only status codes, because an application returning a friendly error page with a 200 status will otherwise be recorded as a fast success.

Trap to avoid

Not validating response content. A load test where the application is failing gracefully reports excellent response times, and this is one of the most common ways a performance test gives a falsely reassuring result.

Q25Mid-levelEnvironment & data

How closely should the test environment match production?

What they are assessing

Realism about a constraint everyone faces.

Model answer

As closely as possible, and it rarely is. What matters most, in order, is architectural equivalence, meaning the same components and topology even if smaller, then data volume, then relative resource ratios, then absolute capacity. A scaled-down environment can still give valid comparative results and find most bottlenecks, provided the scaling is understood and stated. What invalidates results entirely is a structural difference: a single application server where production has a load-balanced cluster, or no caching layer, since the bottleneck you find may not exist in production and the one that matters will be invisible.

Likely follow-up

How do you extrapolate from a half-sized environment?

Q26SeniorEnvironment & data

Why does test data volume matter as much as it does?

What they are assessing

A frequently underestimated factor.

Model answer

Because query behaviour changes with volume in ways that are not proportional. A query without an index is instant on ten thousand rows and catastrophic on ten million, so a test against a small database finds nothing. Query plans also change as the optimiser sees different statistics, so the plan in test may differ from production entirely. Beyond queries, data volume affects cache hit rates, index depth, backup windows and disk behaviour. A performance test against an empty database is one of the most common ways to produce results that look excellent and predict nothing.

Q27SeniorEnvironment & data

Would you run a performance test against production?

What they are assessing

A nuanced question with a defensible answer either way.

Model answer

It is increasingly done and it can be the right call, because it is the only environment that is genuinely representative. The conditions I would require are: explicit sign-off from the business and operations, a window agreed in advance, a kill switch and someone watching who can stop it instantly, careful data handling so test records are identifiable and removable, and ideally shadow traffic or a canary approach rather than full load. What I would not do is run it unannounced, or run a stress test to failure against a live system, because the point of stress testing is to break something and that choice is not a tester call to make alone.

Likely follow-up

How would you mark test data so it can be identified and removed later?

Q28Mid-levelEnvironment & data

What is the risk of running load generators on the same machine as the application?

What they are assessing

A basic but frequently made mistake.

Model answer

The generator competes with the application for CPU, memory and network, so you measure a system under combined load and attribute all of it to the application. It also caps the load you can generate. More subtly, it hides network behaviour entirely, since requests never traverse a real network. The generator should be separate, monitored for its own resource usage, and verified not to be the bottleneck, because a saturated load generator produces a throughput ceiling that is very easily mistaken for an application limit.

Q29Mid-levelExecution

What is a ramp-up period and why does it matter?

What they are assessing

Test design detail with real consequences.

Model answer

The time over which virtual users are added rather than starting all at once. It matters for two reasons. First, realism: real load builds gradually, and an instantaneous thousand-user start is a spike test regardless of what you called it. Second, measurement: during ramp-up the system is warming caches, filling connection pools and compiling code, so results from that window are not representative and should be excluded from the analysis. A gradual ramp also lets you see at what load level degradation begins, which a step-change start cannot show.

Likely follow-up

Why should the steady-state period be measured separately from ramp-up?

Q30Mid-levelExecution

How long should a performance test run?

What they are assessing

Practical judgement.

Model answer

Long enough to reach and hold a steady state, which usually means at least thirty minutes to an hour of measurement after ramp-up for a standard load test. Shorter runs are dominated by warm-up effects and miss anything cyclical such as garbage collection, cache expiry or scheduled jobs. Soak tests run far longer by definition. The test I apply is whether the metrics have stabilised: if response times are still trending at the end of the run, it has not run long enough to tell you anything conclusive.

Q31SeniorExecution

What would you check before accepting a test run as valid?

What they are assessing

Rigour, and whether they question their own results.

Model answer

Error rate first, because a run with significant errors is measuring a failing system and the response times are meaningless. Then whether the load generators themselves were saturated. Whether the applied load actually matched the planned workload, since pacing and failures can silently reduce it. Whether the environment was exclusively used, as another team deploying mid-run invalidates everything. Whether data ran out, causing users to fail or reuse records. And whether the steady-state window is genuinely steady. I would rather discard a run and repeat it than report numbers I cannot defend, because a single bad result published to stakeholders takes months to correct.

Trap to avoid

Reporting results without checking the error rate. Low response times with a thirty percent error rate is the most common false positive in performance testing.

Q32SeniorExecution

How do you ensure results are repeatable?

What they are assessing

Experimental method.

Model answer

By controlling what varies: the same build, the same environment with no concurrent activity, the same data state restored before each run, the same workload model, and the same ramp and duration. Then by repeating runs, since a single execution tells you nothing about variance, and reporting the spread rather than one number. Where results differ between identical runs, that variance is itself the finding and needs explaining before any comparison between releases is credible. Documenting the conditions is what makes a result reproducible by someone else, which is the real test.

Q33Mid-levelResults analysis

What metrics do you look at first after a run?

What they are assessing

Analysis method.

Model answer

Error rate, to establish whether the run is valid at all. Then throughput against the target, to confirm the intended load was actually applied. Then response times by transaction with percentiles, to see what the user experienced. Then the shape over time, because an average across the run hides degradation. Then resource utilisation on each tier: CPU, memory, disk and network on application and database servers. The order matters because each step can invalidate the next: there is no point analysing response times from a run that applied half the intended load.

Q34SeniorResults analysis

Response times are rising steadily through the run while load is constant. What does that suggest?

What they are assessing

Diagnostic reasoning, a very common interview scenario.

Model answer

Something is accumulating. The usual candidates are a memory leak driving increasing garbage collection, a connection or thread pool filling up, unbounded growth in a cache or session store, database bloat from the data the test itself is inserting, or log files filling a disk. The diagnostic step is to correlate the response time curve against memory and garbage collection on the application tier and against database size and lock waits, and to check whether the rise continues after load is removed, which distinguishes an accumulation problem from simple saturation.

Likely follow-up

How would you distinguish a memory leak from connection pool exhaustion?

Q35SeniorResults analysis

Throughput has plateaued but response time keeps climbing as you add users. What is happening?

What they are assessing

The classic saturation signature.

Model answer

The system has reached its maximum throughput and additional users are queueing. This is the knee of the curve, and it is the single most important thing a stress test identifies, because it defines capacity. Beyond that point every extra user adds waiting time without adding work completed. The next question is which resource is saturated, and that is answered from the utilisation data: if a CPU is at a hundred percent it is straightforward, and if nothing is saturated the constraint is a lock, a pool limit, a thread count or a single-threaded component.

Trap to avoid

Reading flat throughput as stability. Flat throughput with rising response time is saturation, and it is the opposite of a good result.

Q36SeniorResults analysis

What is Little’s Law and how do you use it?

What they are assessing

Whether the candidate has any theoretical grounding.

Model answer

Concurrency equals throughput multiplied by response time. It is useful as a sanity check on a test: if the tool reports a hundred concurrent users, throughput of fifty requests per second and an average response time of four seconds, those numbers are inconsistent and one of them is wrong. It is also useful for deriving a target: if the requirement is six hundred transactions per minute with a two-second response time, that implies twenty concurrent in-flight requests, which tells you whether your thread pool configuration is even capable of it.

Q37Mid-levelResults analysis

What does a high standard deviation in response times indicate?

What they are assessing

Reading beyond the headline number.

Model answer

Inconsistency, which users experience as unpredictability and dislike more than consistent slowness. The usual causes are intermittent contention, garbage collection pauses, a cache with a poor hit rate so some requests are served fresh and others computed, uneven load balancing across nodes, or a subset of requests hitting a slow path such as a different query plan for some data. It is a prompt to look at the distribution and the time series rather than to report the mean, since the mean is precisely the statistic that conceals it.

Q38Mid-levelBottlenecks

What are the common bottlenecks in a web application?

What they are assessing

Breadth of diagnostic knowledge.

Model answer

Database most often: missing indexes, inefficient queries, lock contention, or an undersized connection pool. Then application-tier issues: thread pool limits, synchronisation and locking, memory pressure driving garbage collection, and inefficient algorithms over large collections. Then external dependencies, where a slow third-party call holds a thread for its duration. Then infrastructure: CPU, memory, disk I/O and network bandwidth. And configuration limits, which are often the real answer, since a default pool size or worker count set for development is a very common ceiling.

Likely follow-up

Why is the database the most frequent culprit?

Q39SeniorBottlenecks

How do you identify which tier is the bottleneck?

What they are assessing

Systematic diagnosis.

Model answer

By correlating timing across tiers for the same period. If application server CPU is low while response times are high, the application is waiting on something: database, external call or lock. Breaking the transaction down by component, which is what APM does well, shows where the time is spent rather than inferring it. Resource utilisation on each tier identifies saturation. If no resource is saturated anywhere and throughput is capped, the constraint is logical rather than physical, meaning a pool, a thread limit, a lock or a serialised section, and those are found through thread dumps and database wait statistics rather than through utilisation graphs.

Q40SeniorBottlenecks

CPU is at 40 percent, memory is fine, and throughput will not increase. Where do you look?

What they are assessing

Reasoning when the obvious indicators are clean.

Model answer

This is a logical constraint, not a resource one. I would check thread pool and connection pool sizes first, since a pool capped at fifty will hold throughput flat with idle CPU and is the most common answer. Then thread dumps for blocked threads, which reveals lock contention and serialised sections. Then database wait events and lock waits. Then whether a downstream dependency is the actual limit, since threads waiting on a slow external call show as idle CPU. And finally whether the load generator or the network is the constraint rather than the application, which is embarrassing to discover late and worth ruling out early.

Likely follow-up

What would a thread dump showing fifty threads blocked on one monitor tell you?

Q41SeniorBottlenecks

How do you investigate a database bottleneck?

What they are assessing

Depth in the most common problem area.

Model answer

Start from the database rather than from the application. The slow query log or equivalent gives the expensive statements, and execution plans show whether they are scanning when they should be seeking. Wait statistics show whether time is spent on I/O, locks or CPU, which points in completely different directions. Lock and blocking reports identify contention. Connection pool metrics show whether requests are queueing before they even reach the database. The frequent finding is not one slow query but a fast query executed thousands of times per transaction, which an execution count ordering reveals and a duration ordering hides.

Trap to avoid

Only looking for slow queries. The N+1 pattern, where a query taking two milliseconds runs four hundred times per page, never appears in a slow query log.

Q42Mid-levelMonitoring

What server-side metrics do you monitor during a test?

What they are assessing

Practical instrumentation.

Model answer

On every tier: CPU including breakdown by user, system and wait, memory including swap, disk I/O and queue depth, and network throughput and errors. On the application: heap usage, garbage collection frequency and pause duration, thread counts and states, and connection pool utilisation. On the database: active sessions, locks and blocking, buffer cache hit ratio, slow queries and wait events. On the web tier: queued requests and worker utilisation. Without these, a test tells you something is slow and nothing about why.

Q43SeniorMonitoring

What is APM and how does it change performance testing?

What they are assessing

Modern tooling awareness.

Model answer

Application Performance Monitoring tools such as Dynatrace, AppDynamics, New Relic or open-source equivalents instrument the application to trace individual transactions through every tier and method. It changes the work substantially: instead of inferring where time is spent from external measurements, you see the breakdown directly, down to the specific method or query. It turns a finding of checkout is slow into checkout spends 1.8 seconds in a single repository call executing 340 queries, which is actionable rather than a prompt for further investigation. Where APM exists, the performance engineer spends far less time diagnosing and more time on workload design and interpretation.

Likely follow-up

What does APM not give you that a load test does?

Q44SeniorMonitoring

What is garbage collection and why does it matter in performance testing?

What they are assessing

Depth on JVM and similar runtimes.

Model answer

Automatic memory reclamation in managed runtimes. It matters because collection pauses stop application threads, appearing as periodic latency spikes that correlate with nothing in the workload and are frequently the explanation for a bad 99th percentile alongside a fine median. Frequent full collections with little memory recovered indicate a leak or an undersized heap. The metrics worth collecting are pause duration, frequency and the proportion of time spent collecting, and the common misdiagnosis is to treat the resulting latency spikes as a network or database problem because the request timing offers no other clue.

Q45Mid-levelScalability

What is the difference between vertical and horizontal scaling, from a testing perspective?

What they are assessing

Capacity testing design.

Model answer

Vertical adds resources to a machine, horizontal adds machines. For testing, vertical scaling is simpler to measure but hits limits and does not address single points of failure. Horizontal scaling needs to be tested for the things that only appear with multiple nodes: session affinity, distributed cache coherence, load balancer behaviour, and whether any shared resource becomes the new bottleneck. The key measurement is the scaling efficiency, since a system gaining sixty percent throughput from a doubled tier has a shared constraint that will cap it regardless of how many nodes are added.

Q46SeniorScalability

How does auto-scaling affect performance testing?

What they are assessing

Cloud-era considerations.

Model answer

It changes what you are testing. A gradual ramp lets auto-scaling keep up, so the test validates the steady state but not the scaling behaviour. A spike test is what exercises it, and the interesting measurements become scaling latency, meaning how long between the trigger and capacity being available, and what users experience during that gap, which is typically where the failures occur. You also need to test scale-down, since aggressive policies can remove capacity while load is still present. And cost becomes a result in its own right: a system that meets its targets by scaling to forty instances has passed the performance test and possibly failed the business case.

Likely follow-up

What would you measure during the window between a spike and new capacity coming online?

Q47SeniorScalability

What is a circuit breaker and why does it matter in a load test?

What they are assessing

Resilience awareness.

Model answer

A pattern where repeated failures to a dependency cause calls to be short-circuited rather than attempted, protecting the caller from a slow or failing downstream and giving it time to recover. It matters in a load test because it changes behaviour under stress in ways that look like success: once a breaker opens, response times drop sharply because requests are failing fast rather than succeeding. A naive reading of the results shows improved latency at high load, when what has actually happened is that the system has stopped doing the work. This is one of the clearest reasons to validate response content rather than only timing.

Q48Mid-levelTooling

Which performance testing tools have you used and how do they differ?

What they are assessing

Breadth, and the ability to compare rather than list.

Model answer

JMeter is the common open-source choice, GUI-driven for authoring with a large plugin ecosystem, thread-per-user so resource-hungry at high scale. Gatling uses an asynchronous model with a Scala or Java DSL, so it generates far more load per machine and the scripts are code. k6 is JavaScript-based, lightweight and built for CI integration. LoadRunner is the long-standing commercial option with the broadest protocol support beyond HTTP, which still matters for enterprise systems using SAP, Citrix or mainframe protocols. Locust is Python-based and favoured where the team already writes Python.

Likely follow-up

Why does a thread-per-user model limit scale?

Q49SeniorTooling

How would you choose a performance testing tool for a new project?

What they are assessing

Structured selection.

Model answer

Protocol support first, because it eliminates options outright: if the system uses something other than HTTP, most tools are immediately unsuitable. Then the required load and the hardware available to generate it, since that decides whether a thread-per-user model is viable. Then team skills, because a tool nobody can script in produces one engineer who becomes a bottleneck. Then CI integration, if performance testing is to be automated rather than a periodic event. Then licensing cost against budget. I would resist choosing on familiarity alone, but I would weight it heavily, since an unfamiliar tool costs months of productivity on a short engagement.

Q50SeniorTooling

How do you integrate performance testing into CI?

What they are assessing

Modern practice.

Model answer

By running a short, targeted test on a schedule or per build rather than the full suite: a scaled-down load against key transactions, with thresholds that fail the build on regression against a baseline. The important design decisions are what the thresholds measure, which should be percentiles rather than averages, and how much variance is tolerated, since a threshold set too tight produces constant false failures and gets disabled within a month. Full-scale tests remain a scheduled activity against a proper environment, because CI environments are too small and too noisy to produce absolute numbers worth acting on.

Trap to avoid

Promising full-scale load testing in CI. The infrastructure cost and run duration make it impractical for most organisations, and claiming otherwise suggests limited experience of doing it.

Q51Mid-levelReporting

What goes into a performance test report?

What they are assessing

Communication.

Model answer

The objective and the requirements being tested. The workload model and the assumptions behind it. The environment, including how it differs from production. The results against each requirement, with percentiles rather than averages. Observations and bottlenecks identified, with supporting evidence. A clear conclusion and recommendation. And the limitations, stating what was not tested and what cannot be concluded. The limitations section is the one most often omitted and the most important, because without it readers assume the test covered everything.

Q52SeniorReporting

How do you present results to a non-technical stakeholder?

What they are assessing

The skill that distinguishes senior performance engineers.

Model answer

In terms of user experience and business risk rather than metrics. Not the 95th percentile of the search transaction is 4.2 seconds, but one in twenty searches takes over four seconds at peak, which in our traffic is about three hundred customers an hour, and we know users abandon at around three. Then the cause and the options with their costs. A graph of response time against user count is usually the one chart worth showing, because the knee is visually obvious and makes capacity intuitive. The failure mode is presenting a dashboard of thirty metrics and letting the stakeholder guess which matter.

Likely follow-up

How do you report when the result is that the system meets its target but only just?

Q53SeniorReporting

The system fails its performance requirement two days before release. What do you do?

What they are assessing

Behaviour under pressure.

Model answer

Establish the facts quickly and precisely: which transactions, at what load, by how much, and whether it is a regression against a previous release or a pre-existing condition that was never tested. Verify the result is valid rather than an artefact of the environment, because a false alarm at this point is expensive. Then quantify the user impact in business terms, and present the options: ship with a known limitation and a mitigation such as a capacity increase or a feature flag, delay, or reduce scope. My job is to make the risk clear and the decision informed, not to make the decision. What I would not do is present it as a binary pass or fail, since the useful statement is almost always that the system supports X when the target was Y.

Q54LeadReporting

How do you justify investment in performance testing to management?

What they are assessing

Commercial framing.

Model answer

In the terms they already care about. The cost of an outage or a degraded peak period, using the organisation’s own revenue-per-hour figures where available. The documented relationship between latency and conversion, which is well established and large enough to make the case on its own for a transactional business. The cost of over-provisioning, since many organisations spend substantially on infrastructure to compensate for problems a few days of tuning would fix, and that is a line item the finance team can see. And the cost of discovering it late, where a fix at design time is an architecture decision and the same fix after launch is an emergency. The argument that fails is testing is important as a principle.

Q55Mid-levelFundamentals

What is the difference between client-side and server-side performance?

What they are assessing

Scope awareness.

Model answer

Server-side performance is how quickly the backend produces a response, which is what protocol-level load tests measure. Client-side performance is everything after that: downloading assets, parsing and executing JavaScript, rendering, and the sequencing of further requests the page triggers. Users experience the total. A system can have a 200 millisecond server response and a six-second perceived load because of unoptimised images, render-blocking scripts and a large JavaScript bundle. They need different tools and different fixes, and reporting only server-side numbers as performance is a common and significant omission.

Q56SeniorFundamentals

What are Core Web Vitals and are they relevant to performance testing?

What they are assessing

Current front-end performance awareness.

Model answer

Google’s user-centric metrics: Largest Contentful Paint for loading, Interaction to Next Paint for responsiveness, which replaced First Input Delay, and Cumulative Layout Shift for visual stability. They are relevant because they are what users and search ranking actually respond to, and because they measure experience rather than server timing. They are measured with Lighthouse, WebPageTest or real user monitoring rather than with a load testing tool, so they sit alongside load testing rather than within it. The connection worth making is that they should be measured while the system is under load, since a page that scores well on an idle system can score badly at peak.

Q57Mid-levelWorkload modelling

How do you model a peak such as a sale or a ticket release?

What they are assessing

Modelling an atypical but high-stakes pattern.

Model answer

Not as a scaled-up version of normal traffic, because the shape is completely different. The arrival pattern is near-instantaneous rather than ramped, the transaction mix is narrow with everyone doing the same thing, caching behaves differently because all users want the same records, and contention concentrates on a small number of rows. So the model needs a spike arrival profile, a mix weighted heavily to the specific journey, and deliberate contention on the same inventory. Modelling it as ten times the usual mix produces a test that passes and a launch that fails.

Likely follow-up

Why does everyone wanting the same item change the bottleneck?

Q58SeniorScripting

How do you handle dynamic content such as a shopping basket during a load test?

What they are assessing

State management in scripts.

Model answer

By correlating the identifiers the application returns and carrying them through the flow, treating each virtual user as having its own independent state. The design questions are whether users share a data pool and what happens when two pick the same item, whether the basket is cleared between iterations or accumulates, and whether the data created is cleaned up afterwards. Accumulating state is a frequent cause of a soak test degrading, where after six hours every virtual user has a basket with three thousand items and the slow response times are an artefact of the test rather than a finding.

Q59SeniorEnvironment & data

How do you reset state between test runs?

What they are assessing

Repeatability in practice.

Model answer

A database restore from a known snapshot is the cleanest and is worth the time it takes, because it guarantees the same starting point. Where that is impractical, targeted cleanup scripts removing the records the test created, combined with a naming or flagging convention that makes them identifiable. Caches need clearing or warming deliberately, and whichever you choose must be consistent, since comparing a cold-cache run with a warm-cache one produces a difference that has nothing to do with the change being tested. Application restarts reset pools and heap, and whether to do that between runs is a decision to make once and keep.

Q60Mid-levelExecution

What is distributed load generation and when do you need it?

What they are assessing

Scaling the test itself.

Model answer

Running load generators across several machines coordinated by a controller, needed when one machine cannot produce the required load without saturating its own CPU, memory, network or available ports. The signs you need it are generator CPU above about seventy percent, or port exhaustion, which appears as connection errors that look like application failures. It also allows generating load from several geographic locations, which matters when network latency is part of what you are measuring. The thing to watch is that the generators themselves are monitored, since a saturated generator silently caps throughput.

Q61SeniorResults analysis

How do you compare two releases for performance regression?

What they are assessing

Comparative method.

Model answer

Same environment, same data state, same workload, same duration, ideally run close together to avoid environmental drift, and repeated at least twice each so variance is known. Then compare at the transaction level with percentiles rather than overall averages, because a regression in one transaction can be invisible in an aggregate. The judgement is whether an observed difference exceeds normal run-to-run variance, which is why knowing that variance matters: declaring a five percent regression on a system that varies by eight percent between identical runs is noise reported as a finding.

Likely follow-up

How many runs would you want before declaring a regression?

Q62Mid-levelBottlenecks

What is the N+1 query problem?

What they are assessing

A specific and very common performance defect.

Model answer

Where retrieving a list of N items triggers one query for the list and then one additional query per item, usually because an ORM is lazily loading a relationship inside a loop. It is invisible in development with ten records and crippling with a thousand. It does not appear in slow query logs because each individual query is fast. It is found by looking at query counts per transaction rather than query duration, which APM shows directly, and the fix is eager loading or a join. It is worth naming specifically because it is probably the single most common application-level performance defect in web systems.

Q63SeniorMonitoring

What is the difference between synthetic monitoring and real user monitoring?

What they are assessing

Production observability.

Model answer

Synthetic monitoring runs scripted transactions on a schedule from known locations, giving consistent, comparable measurements and detecting outages even with no traffic, but it only measures the journeys you scripted. Real user monitoring instruments actual sessions, so it reflects the full diversity of devices, networks and journeys, including the slow ones you would never have scripted, but it is noisy and gives no data where there are no users. They complement each other, and for a performance engineer RUM is also the best available source for workload modelling, since it describes what users actually do rather than what we assume.

Q64SeniorScalability

What is the difference between performance and capacity testing?

What they are assessing

A distinction that affects how a test is designed.

Model answer

Performance testing asks whether the system meets its targets at a defined load. Capacity testing asks how much load the system can support while still meeting them, which is a different question requiring a different test: an incremental ramp through increasing load levels, measuring at each step until targets are breached. The output of capacity testing is a number used for planning, such as the system supports 1,400 concurrent users within SLA on this configuration. It is the more useful of the two for infrastructure decisions, and the less commonly done.

Q65Mid-levelRequirements

What performance requirements would you set for a mobile application?

What they are assessing

Context-specific thinking.

Model answer

Different from web, because the constraints are different. App launch time, both cold and warm. Screen transition responsiveness. Behaviour on a poor network, which matters far more than on desktop, so targets should be stated for 3G and intermittent connectivity rather than only for ideal conditions. Battery and data consumption over a session. Memory footprint, since the operating system terminates heavy applications in the background. And API response times under the same conditions. Testing only on a good office network and a current flagship device misses most of what users actually experience.

Q66SeniorTooling

How would you performance test an API rather than a web application?

What they are assessing

A very common modern scenario.

Model answer

It is simpler in some ways and harder in others. Simpler because there is no browser rendering and correlation is usually cleaner. Harder because the workload model must reflect how clients actually call it, which for a public API may be unknown and highly variable, and because a single client behaving badly can dominate. The things to test beyond throughput and latency are rate limiting behaviour, payload size effects since a large response changes everything, pagination performance at depth, and authentication overhead where token validation involves a remote call. Error responses under load matter too, since an API that returns 500s quickly can look fast.

Q67SeniorExecution

How do you performance test a system with third-party dependencies?

What they are assessing

A practical and contractual problem.

Model answer

Carefully, because load testing someone else’s system without agreement is at best a breach of terms and at worst an attack. The options are to obtain explicit permission and a window, to use the provider’s sandbox if it is representative, which it usually is not, or to stub the dependency with a service virtualisation layer configured to match its real latency and error behaviour. Stubbing is usually right, and the critical part is that the stub must reproduce realistic delay, because replacing a 400 millisecond external call with an instant mock removes the thread-holding behaviour that was the whole point of including it.

Trap to avoid

Load testing a third-party API without permission. It is the kind of mistake that ends an engagement, and interviewers at senior level sometimes ask specifically to see whether you raise it.

Q68LeadReporting

How do you set up a performance testing practice from scratch?

What they are assessing

Building capability rather than running tests.

Model answer

I would start with one system that matters and produce a result people act on, because credibility comes from a finding that prevented something, not from a strategy document. In parallel, the foundations: an environment that is representative enough to be believed, a workload model derived from production data, baselines so later results mean something, and monitoring, since without it every test ends at it is slow. Then tooling and scripts in version control, then CI integration for regression detection. The organisational work matters as much: getting performance requirements written at the same time as functional ones, which is what moves the practice from reactive to preventive.

Likely follow-up

What would you deliberately not do in the first three months?

Q69SeniorResults analysis

What is coordinated omission and why does it matter?

What they are assessing

A genuinely advanced topic that distinguishes serious practitioners.

Model answer

A measurement error where a load generator, having sent a request, waits for the response before sending the next, so when the system stalls the generator stops issuing requests. The requests that would have been slow are never sent, and the recorded results omit exactly the worst cases, understating high percentiles substantially. It matters because it means a tool can report an excellent 99th percentile for a system that periodically freezes. The mitigations are to use a constant arrival rate rather than closed-loop pacing, which several modern tools support, and to be aware that any closed-loop test under-reports tail latency.

Q70Mid-levelMonitoring

What is the difference between utilisation, saturation and errors?

What they are assessing

A useful diagnostic framework.

Model answer

This is the USE method. Utilisation is the proportion of time a resource is busy. Saturation is the degree to which work is queued waiting for it. Errors are failure counts. The reason all three are needed is that utilisation alone is misleading: a CPU at seventy percent with a long run queue is saturated and a disk at ninety percent with no queue may be fine. Saturation is usually the better early indicator, since queueing begins before utilisation reaches its ceiling, and it is the metric most often not collected.

Q71SeniorWorkload modelling

How do you validate that your workload model is realistic?

What they are assessing

Self-checking, which few candidates mention.

Model answer

By comparing the test against production on the dimensions you can observe. The transaction mix should match production ratios. The resulting resource profile on the servers should resemble production at a comparable load, so if production runs database-heavy and the test runs CPU-heavy on the application tier, the model is wrong somewhere. Throughput per user should be plausible. And where production data exists, a test at production load should reproduce production response times, which is the strongest validation available and the one worth doing before trusting any result at higher loads.

Q72Mid-levelFundamentals

What is a bottleneck, and can you remove all of them?

What they are assessing

Conceptual clarity.

Model answer

The resource or component that limits overall throughput. You cannot remove all of them, because there is always one: fixing the current constraint moves it elsewhere. That is the correct outcome rather than a failure, and the practical question is whether the new limit is high enough for the requirement. The implication for testing is that performance work is iterative by nature, and a plan that assumes one round of testing and one fix is unrealistic. It also means the right stopping point is when the requirement is met with margin, not when no bottleneck remains.

Q73SeniorEnvironment & data

What is service virtualisation and when would you use it?

What they are assessing

A technique for dependency problems.

Model answer

Simulating a dependency with a configurable stand-in that reproduces its interface, latency and error behaviour. It is used when the real dependency is unavailable, not under your control, rate-limited, expensive per call, or when you need it to behave in ways it will not on demand, such as responding slowly or failing. For performance testing specifically, the critical configuration is latency, since a virtualised service that responds instantly changes the concurrency behaviour of the system under test completely and will understate thread pool pressure.

Q74LeadFundamentals

How do you decide when performance testing is enough?

What they are assessing

Judgement about stopping, which is rarely discussed.

Model answer

When the remaining risk is understood and acceptable to the people who own it, which is a business judgement I inform rather than make. Practically that means the stated requirements are met with margin on a representative environment, the capacity limit is known so there is a planning figure, failure behaviour under overload has been observed rather than assumed, and the limitations of what was tested are documented. It is never complete in an absolute sense. The engineer’s obligation is to be clear about what has not been covered, because the risk of stopping is only acceptable if it is visible.

Q75SeniorBottlenecks

What is thread pool exhaustion and how does it present?

What they are assessing

A specific failure mode.

Model answer

When all worker threads are occupied, so new requests queue and eventually time out or are rejected. It presents as throughput capped well below resource limits, with CPU and memory comfortable, and response times rising sharply at a particular load as queueing begins. The usual underlying cause is threads being held by slow operations rather than too few threads: a two-second external call holds a thread for two seconds, so a pool of two hundred supports only a hundred such requests per second. Increasing the pool size often makes it worse by adding contention, which is why the real fix is usually to make the slow operation faster or asynchronous.

Likely follow-up

Why would increasing the thread pool size sometimes reduce throughput?

Q76Mid-levelResults analysis

What error rate is acceptable in a performance test?

What they are assessing

Standards.

Model answer

For a load test validating expected conditions, close to zero, because errors at expected load are a functional failure rather than a performance characteristic. A commonly used threshold is under one percent, though I would want to know what the errors are rather than accept a percentage: a small number of timeouts means something different from a small number of 500s. In a stress test beyond the breaking point, errors are the expected result and the question becomes how the system fails, whether it degrades gracefully, and whether it recovers.

Q77SeniorScalability

What does graceful degradation mean, and how would you test it?

What they are assessing

Resilience thinking.

Model answer

That as load exceeds capacity the system continues serving a reduced service rather than failing entirely: shedding non-essential features, queueing with honest feedback, serving cached or stale content, or rejecting some requests quickly rather than accepting all and timing out. Testing it means deliberately exceeding capacity and observing what happens to users, not just whether the servers survive. The specific things to look for are whether failures are fast or slow, since slow failures hold resources and cascade, whether errors are informative, and whether the system recovers on its own once load returns to normal.

Q78SeniorTooling

How do you performance test asynchronous or event-driven systems?

What they are assessing

A scenario where request-response assumptions break down.

Model answer

The response time of the initial call is close to meaningless, because the work happens afterwards. What matters is end-to-end latency from event submission to the effect being observable, queue depth and whether it grows without bound, consumer lag, and processing throughput against arrival rate. So the test has to correlate a submitted message with its eventual outcome, usually by instrumenting both ends and matching on an identifier. The key failure to look for is the queue growing steadily, which means consumption is slower than production and the system will fail eventually regardless of how good the latency looks at the start.

Likely follow-up

What does a steadily growing queue depth tell you that latency does not?

Q79LeadReporting

A developer disputes your findings, saying the test is unrealistic. How do you respond?

What they are assessing

Handling challenge, which performance engineers face constantly.

Model answer

By taking it seriously, because quite often they are partly right and the workload model is where I would look first myself. I would ask specifically what is unrealistic, since a general objection is usually discomfort while a specific one is information. If the point is valid, I change the model and rerun, which strengthens the finding rather than weakening it. If it is not, I show the basis for the model, ideally production data, which moves the conversation from opinion to evidence. What I avoid is defending the test as a matter of authority, because the goal is a system that performs, not a report that survives review.

Q80Mid-levelExecution

What is a smoke test in performance terms?

What they are assessing

Practical process.

Model answer

A short run with a very small number of users, typically one to five, to validate that the script works, correlation is holding, data is being consumed correctly and the environment is responding, before committing to a full run. It costs minutes and saves hours, because the most common cause of a wasted two-hour test is a script failing at step four for every user. It also gives a single-user baseline, which is genuinely useful: the response time with no contention is the floor, and comparing it to the loaded result separates inherent slowness from contention.

Q81SeniorRequirements

How do you handle a requirement that the system support an unrealistic number of users?

What they are assessing

Professional pushback.

Model answer

By testing it anyway and reporting what it costs, rather than arguing about the number. If the stated requirement is ten thousand concurrent users and the realistic figure is eight hundred, the useful output is a capacity curve showing what the system supports at each infrastructure level and what each step costs. That converts an argument about a number into a business decision about spend. Where I would push back is on the derivation, by showing how the figure was reached and asking how theirs was, since an invented requirement usually comes from a misunderstanding such as treating registered users as concurrent ones.

Q82LeadFundamentals

What is the most common mistake you see in performance testing?

What they are assessing

A closing question that reveals depth of experience.

Model answer

Running a technically competent test against a workload model nobody validated. Everything downstream depends on it, and it is the part that gets the least scrutiny, because the tooling is visible and the assumptions are not. Close behind are reporting averages instead of percentiles, which hides the user experience that matters, and not checking the error rate, which produces excellent-looking results from a failing system. All three share the same character: they do not produce an obviously wrong answer, they produce a confident one, which is considerably more dangerous than a test that visibly fails.

Where Interviews Are Won

What performance interviews actually separate on

Defining load against stress takes two minutes. These four areas decide the outcome, and all four come from having run a real test and defended the result.

The workload model

Everything downstream depends on it and it gets the least scrutiny. Expect to be asked how you derive load when the business cannot tell you.

Percentiles over averages

An average of 1.2 seconds is consistent with ten percent of users waiting nine. Candidates who report only means lose this round quickly.

Reading saturation

Flat throughput with climbing response time is the knee of the curve, not stability. Mistaking it for a good result is a common and costly error.

Validating your own run

Excellent response times with a thirty percent error rate is the classic false positive. Checking errors before reporting marks out an engineer.

Who Wrote This

Written by engineers who run these tests

This bank was written and reviewed by QAble performance engineers who model, run and interpret load tests on client systems, including the ones where the finding was unwelcome: workload models built on a misunderstanding of concurrent users, soak tests that degraded because the test data accumulated rather than the system leaking, and reports of excellent latency produced by a circuit breaker that had quietly opened.

Answers are pitched at the level marked on each question, and the tooling questions stay deliberately neutral. JMeter specifics live in the JMeter bank, so nothing here is padding. If you think an answer here is wrong, we would genuinely like to hear it.

Tell us what we got wrong

Unsure what your system actually holds?

QAble runs performance engineering end to end: workload modelling from production data, execution against a representative environment, and bottleneck diagnosis rather than a report.

Performance testing services

More question banks

View all

UFT interview questions

Question bank
82 UFT One questions across the object repository and identification, Smart Identification, descriptive programming, checkpoints, actions, recovery scenarios and framework design.

Test manager interview questions

Question bank
82 questions at organisational level: QA strategy, operating model, budgets and cost of quality, resourcing, vendor selection, tooling, metrics, governance and transformation.

Test lead interview questions

Question bank
82 questions at team and delivery level: planning and estimation, allocation, risk-based strategy, triage, reporting, stakeholder management, mentoring and the difficult conversations.

SQL for testers interview questions

Question bank
82 questions on query skill for QA roles: joins and the anti-join pattern, aggregation, the NOT IN null trap, set operations for reconciliation, window functions and safe data modification.

Mobile testing interview questions

Question bank
82 questions across device strategy, platform differences, upgrade paths, interrupts and the app lifecycle, network conditions, performance, mobile security and staged release.

REST Assured interview questions

Question bank
82 questions across the Java DSL, GPath body assertions, schema validation, serialisation with POJOs, authentication, reusable specifications, filters and parallel execution thread safety.

BDD interview questions

Question bank
82 questions on the practice rather than the tooling: discovery, formulation and automation, example mapping, declarative scenario design, living documentation and the anti-patterns.

JUnit interview questions

Question bank
82 JUnit 5 questions across the three module architecture, annotations and lifecycle, assertThrows and assertAll, parameterized tests, the extension model, Mockito and migration from JUnit 4.

Banking domain testing interview questions

Question bank
82 domain questions across the general ledger and double entry, payments and reversals, cards, lending and interest, KYC and AML, the batch and end of day cycle and reconciliation.

Agile testing interview questions

Question bank
82 questions across testing inside the sprint, the agile testing quadrants, user stories and acceptance criteria, definition of done, automation, regression strategy and the anti-patterns.

Maven interview questions

Question bank
82 questions for automation roles across the POM and lifecycle, dependency scopes and transitive conflicts, Surefire and Failsafe, profiles, multi module builds and CI.

LoadRunner interview questions

Question bank
82 questions across VuGen scripting and the action sections, correlation and parameterisation, transactions and pacing, Controller scenario design, Analysis and LoadRunner Enterprise.

Salesforce testing interview questions

Question bank
82 questions across the order of execution, governor limits and bulkification, the layered sharing model, sandboxes and refresh, Flow, Lightning locators and seasonal release regression.

Functional testing interview questions

Question bank
82 questions across test levels and types, equivalence partitioning and boundary analysis, decision tables, risk based prioritisation, exploratory testing and the automation boundary.

Robot Framework interview questions

Question bank
82 questions across keyword design and abstraction, variables and scope, SeleniumLibrary and the Browser library, custom Python libraries, tags, templates and parallel execution with Pabot.

Jira interview questions

Question bank
82 questions for QA roles across workflows and transitions, JQL, the defect lifecycle, boards and sprints, test management add-ons, reporting and permissions.

Cypress interview questions

Question bank
82 questions across the command queue and retry-ability, selectors, intercept and network stubbing, component testing, CI and parallelisation, and the real limitations.

Software testing interview questions

Question bank
82 questions for freshers through to lead, across fundamentals, the testing lifecycle, test design technique, defect management, agile practice and strategy.

JMeter interview questions

Question bank
82 questions across test plan elements, correlation, timers and pacing, distributed execution, results analysis and troubleshooting.

ETL testing interview questions

Question bank
82 questions across warehouse modelling, slowly changing dimensions, source to target validation, incremental loads and the SQL that verifies them.

TestNG interview questions

Question bank
82 questions across annotations and execution order, data providers and factories, groups, dependencies, parallel execution, listeners and the suite XML.

Tosca interview questions

Question bank
82 questions across modules and scanning, TestCase Design, reusable blocks, buffers and expressions, distributed execution and risk based testing.

Postman interview questions

Question bank
82 questions across variable scopes and precedence, scripting and chaining, assertions and schema validation, authentication, data driven runs and Newman in CI.

Cucumber interview questions

Question bank
82 questions across BDD practice, Gherkin, step definitions and expressions, hooks, tags, data tables, shared state, parallel runs and the anti-patterns.

Database testing interview questions

Question bank
82 questions across schema and constraints, verification SQL, data integrity, transactions and isolation, indexes, migrations, security and NoSQL.

Appium interview questions

Question bank
82 questions across architecture, capabilities, locator strategies, drivers, gestures, hybrid contexts, parallel execution and troubleshooting.

Manual testing interview questions

Question bank
65 questions across fundamentals, test design, defect management, agile, scenarios and lead-level strategy, with model answers and follow-ups.

Selenium interview questions

Question bank
50 questions across WebDriver architecture, locators, waits and flakiness, interactions, framework design, Grid and CI, with model answers and follow-ups.

Playwright interview questions

Question bank
34 questions across architecture, locators, auto-waiting, assertions, fixtures, network mocking, tracing and parallelism.

API testing interview questions

Question bank
42 questions across HTTP semantics, schema validation, authentication, API security, tooling, contract testing and performance.

Automation testing interview questions

Question bank
30 tool-agnostic questions on what to automate, framework design, flakiness, CI/CD, test data, metrics and ROI.

SDET interview questions

Question bank
30 questions across coding, data structures, framework and system design, CI/CD, testability and quality strategy.

Preparing for interviews, or need to know what your system really holds?

QAble runs load, stress and soak testing for BFSI, gaming, healthcare and SaaS platforms with ISTQB-certified engineers. Start with a free QA audit.

Talk to QA Advisor