Cursor Rollouts and Security Review: Tracking Deployments by Environment in 2026
Hello HaWkers, shipping a change does not finish the job: code can pass its tests and still fail when it meets real traffic. On September 23, 2026, Cursor announced Rollouts, which monitors the health of each change after deployment, and Security Review, which looks for exploitable flaws in pull requests. The release brings two familiar questions into the review workflow: did the change work in production, and did it introduce a security weakness?
How does your team answer those questions today? This guide explains what both features do, how to set up useful signals for evaluating them, and where human judgment remains essential. The code examples below show observability practices you can adapt to your service. They are not an official Cursor API.
What Cursor announced and why it matters
The official Cursor announcement introduces two bots for the final stage of software delivery. Rollouts watches a change as it is deployed and reports health by environment. Security Review examines pull requests in the context of the repository and comments on potentially exploitable security problems. According to the announcement, both are available on Teams and Enterprise plans. That is the vendor's published availability information; do not assume either feature has been enabled automatically for your project.
A conventional review examines the diff before the merge. Automated tests help find unexpected behavior in known scenarios. Neither one alone confirms what happens after a change reaches its running environment. A registration path might work in a test and then produce more errors only when a particular combination of real data reaches it. Rollouts tries to connect the deployed commit with the signals that appear afterward.
Security Review addresses a different point in the process: before accepting code that could expose data, bypass authorization, or execute untrusted input. Its purpose is to complement the team's review. The announcement explicitly separates style and code quality, which are assigned to Bugbot, from security findings. That distinction helps reviewers avoid treating a cosmetic suggestion like a concrete attack path.
How Rollouts monitors a change
When a pull request opens, Rollouts reads its diff and the affected systems. It then posts a monitoring plan as a comment. The plan describes perceived risks, the intended effect of the change, the signals it will inspect, and any instrumentation gaps. You can edit the comment, and the bot will use the team's revised version. That matters because the people who know the product should be able to correct a plan that misunderstood the PR's intent.
After the commit deploys, Rollouts wakes on the deployment event and checks logs, metrics, and traces according to the plan. It reports results separately for each environment. The same commit can look healthy in staging and show a regression in production. According to Cursor, the reported outcomes are a healthy change, a detected regression, or an inconclusive result. “Inconclusive” is a valid answer when the application does not emit the necessary signals. It should never be interpreted as approval.
The announcement says Rollouts checks both the intended effect and signals such as errors and latency. Imagine a PR that reduces calls to an external dependency. Lower latency would be useful, but more incorrect responses would wipe out that benefit. The plan needs to capture both sides of the hypothesis. Otherwise, one metric can tell a reassuring but inaccurate story.
If the bot suspects a regression, it identifies the change and alerts its author. Depending on configuration, it can open a revert pull request for review or pass the finding to a cloud agent for a fix. Cursor's documentation also makes clear that Rollouts currently does not merge or roll back on its own. Keep that boundary explicit when assigning operational responsibilities.
First practical step: make the effect observable
A post-deployment analysis tool can only reach conclusions from signals that exist. Before enabling the bot, choose a critical operation and record its outcome consistently. The following example is a minimal Node.js server: it responds at /checkout, measures duration, and writes a JSON event. Run it with node server.mjs, then visit the route to inspect the log. In a real application, remove personal data before recording events and send the output to your telemetry provider.
// server.mjs — educational example of structured logging
import http from 'node:http';
import { performance } from 'node:perf_hooks';
http.createServer((request, response) => {
const started = performance.now();
const isCheckout = request.url === '/checkout';
const status = isCheckout ? 200 : 404;
response.writeHead(status, { 'content-type': 'application/json' });
response.end(JSON.stringify({ ok: isCheckout }));
// Record operational signals only, without customer data.
console.log(JSON.stringify({
route: isCheckout ? '/checkout' : 'unknown',
status,
durationMs: Math.round(performance.now() - started),
environment: process.env.APP_ENV ?? 'local',
commit: process.env.GIT_SHA ?? 'unknown'
}));
}).listen(3000);The commit identifier and environment let you separate events from the new version from earlier events. They do not replace the correlation performed by Rollouts' deployment integration, but they make a human investigation easier. Avoid logging tokens, email addresses, request bodies, or sensitive identifiers merely to “improve observability.” If your instrumentation already produces metrics and traces, use the same names for route, environment, and version across all of them.
Write a plan that can be disproved
A good monitoring plan says what should change and what must not get worse. “Check the logs” is not a criterion: you can read any log without answering the question. For a checkout PR, specify which flow changed, which environments will receive the commit, and which signals would indicate failure. The text below is an example of a human comment you could use to revise the bot's plan, not a format required by Cursor.
Change: reduce repeated calls during checkout.
Expected effect: shorter duration for the /checkout route.
Guardrail signals: 5xx response rate and successful orders.
Segmentation: compare events from the same environment and commit only.
Known gap: there is no metric for abandonment at each step yet.
Decision: if order success data is missing, mark the result inconclusive.Notice the last line. An application might respond faster because it stopped performing an essential step. Without a functional success signal, improved latency alone does not prove health. The same reasoning applies to authentication, payments, and permissions: design the plan to detect the error an aggregate indicator might hide.
In a low-traffic service, a short observation window may not produce enough samples. The right response might be to observe longer or perform a targeted check, not declare victory. In a service spanning several regions, the global average can conceal a local regression. Reviewing the plan gives your team a place to specify the breakdown that actually fits your architecture.
An independent check for the team
Even with automation, it helps to keep a simple rule that the team understands and can run. The example below reads two JSON summaries, calculates error rates, and prints a comparison. It is not Rollouts' algorithm, does not set a universal threshold, and does not decide on a rollback. It makes one of the plan's hypotheses explicit.
// compare.mjs — run: node compare.mjs before.json after.json
import { readFileSync } from 'node:fs';
const read = (path) => JSON.parse(readFileSync(path, 'utf8'));
const [beforePath, afterPath] = process.argv.slice(2);
if (!beforePath || !afterPath) {
throw new Error('Provide both metric files.');
}
const before = read(beforePath);
const after = read(afterPath);
const rate = (sample) => sample.requests === 0
? null
: sample.errors / sample.requests;
console.log(JSON.stringify({
environment: after.environment,
commit: after.commit,
beforeErrorRate: rate(before),
afterErrorRate: rate(after),
beforeLatencyP95Ms: before.latencyP95Ms,
afterLatencyP95Ms: after.latencyP95Ms
}, null, 2));The input files must represent comparable periods and the same operation. If either has zero requests, its rate will be null instead of a misleading division. A rising error rate might have another simultaneous cause; a falling rate might reflect a change in the traffic mix. Use the comparison to start an investigation with traces and incident history before attributing cause to the commit.
What Security Review looks for in a PR
Security Review analyzes code against the rest of the repository and posts a review comment with security findings. According to Cursor, it looks for SQL, command, and template injection; authentication and authorization bypasses; secrets committed to source code; SSRF; unvalidated redirects; unsafe deserialization; and dependency changes that introduce known vulnerabilities. It also tries to follow where user-controlled information enters the program and how that information travels through it.
Each finding includes a severity level, an attack path, and a suggested fix. A reviewer can dismiss a finding with an explanation; in that case, the same issue will not be raised again on that PR. Draft pull requests are ignored, according to the launch page. The moment a PR leaves draft status is therefore part of your review workflow: do not depend on a security comment while the PR remains a draft.
Teams can also apply their own rules. Cursor's examples include restricting which paths may make external calls or which tables a request handler must never query. Such rules express local architectural decisions that a generic vulnerability list does not know. Document them with positive and negative examples so that people and tools can reach the same conclusion.
This small example shows the sort of authorization flaw that deserves attention. It does not represent a guaranteed detection by the product. The function must verify that the resource belongs to the authenticated user before returning it; a lookup by identifier alone could expose another account's data.
// Educational example: authorization tied to the authenticated user.
export async function getInvoice(database, invoiceId, userId) {
if (!userId) throw new Error('Authentication required');
const invoice = await database.findInvoice(invoiceId);
if (!invoice || invoice.ownerId !== userId) {
throw new Error('Invoice unavailable');
}
return invoice;
}To strengthen this protection, keep authorization tests and threat reviews in your process. The bot may point to a suspicious path, but it cannot necessarily know every business rule, contractual exception, or temporary permission in your application. Our guide to web security and the OWASP Top 10 also organizes useful risk categories for a discussion with your team.
How to start without losing operational control
Cursor instructs users to enable the bots in the automations area. For Rollouts, connect version control, the deployment system, and the telemetry provider. The page mentions Origin or GitHub for code and Datadog, among other providers, for signals. After configuration, the feature starts tracking the next pull request. Feature flag integration was announced as a future capability, so do not design an operation that depends on it as if it were already available.
Start with one repository and an important flow that already has dependable telemetry. Review the generated plan, confirm that the commit and environment are correct, and watch a low-risk deployment. Record when the result is healthy, when an alert appears, and when data is missing. That record helps you identify whether the bottleneck is the detector, the instrumentation, or the definition of the expected effect.
For Security Review, choose the repositories to analyze and agree on a triage routine: who reviews findings, who explains dismissals, and who fixes problems before the merge. An automated suggestion without an owner becomes noise. A dismissal without a verifiable reason can hide a real problem. Bring the bot's comment into the same discussion where the team already decides on tests, architecture, and risk.
The announcement mentions usage credits for an initial period and estimates for Teams and Enterprise plans. Because this is a promotional condition that may change, check the official page and your account dashboard before forecasting costs. The adoption decision should consider the value of catching a regression or vulnerability early as well as the time spent investigating inconclusive alerts.
Looking ahead: continuous delivery needs continuous evidence
The most interesting part of this release is how it brings together three moments that many teams handle separately: the PR's intent, behavior after deployment, and exposure to security risks. Rollouts turns intent into an observable plan; Security Review tries to identify exploitable paths before a change goes in. Together, they can shorten the time between a code change and a well-formed question about its effects.
They also expose important limits. An incorrect plan may watch the wrong metric. An environment without useful logs may produce an inconclusive result. An automated finding may need business context. A revert PR still needs human review. I would therefore start with the quality of the signals and clear ownership, then measure whether the automation actually improved the team's deployment routine.
If you ship a change today, try stating in one sentence the expected effect, the signal that proves it, and the signal that would make you stop. If you cannot answer those three points, the problem is visible before any bot gets involved. Cursor offers a new way to put these questions inside a PR; it is up to the team to give them reliable data and a responsible decision.
Let's go! 🦅
📚 Want to Keep Up With What Is Coming?
This article covered Cursor Rollouts and Security Review, but the ecosystem changes every week and not every development becomes an article here.
On X, I share what I am testing, behind-the-scenes notes from my projects, and new tools before they become blog posts.
Follow Me There
💡 Daily content about development, careers, and the tools I actually use

