GPT-6 Astra Found Questions My First Security Review Missed
The most useful thing GPT-6 Astra has done for me so far was not writing code. It was looking at security work I thought had already been reviewed carefully — and giving me a different set of questions to answer.
I started using Astra alongside Fable 5.1 for defined checks in boostN: database RLS policies, API authorization flows and the places where tenant isolation can break. The idea was not to crown a winner. I wanted a second, independent reviewer for the same bounded problem.
What surprised me was the precision. Astra repeatedly surfaced assumptions, paths and edge cases that Fable had not raised in its first pass. Fable’s reviews were often solid. Astra did not make them irrelevant; it made them more useful by challenging them from another angle.
This is not a benchmark, and it is not a claim that Astra will find every vulnerability. It is a field note from work on my own systems. But after several runs, I no longer see the model as an interesting release to watch. For structured security checks, it has become a tool I actively want as a second pair of eyes.
Our most useful learning
1. Astra has been the most useful model I have tested for security work so far. I have used it for RLS policies, API authorization and tenant isolation, but those are only examples. It is equally useful for forming attack hypotheses, inspecting unfamiliar code paths, building small verification tools and turning a vague risk question into a focused, testable check. I am not a cybersecurity specialist, so this is not a claim about the level of a professional audit — it is a practical observation from the work I have done with it.
2. It understands rough real-world briefs. With many models, I still need to explain the real intention behind a task after the first answer. Astra is the first non-Claude model I have used that consistently catches that intention from imperfect input: a rough voice transcription, sentence fragments or a brief with missing context. That makes it useful beyond security work too — for working through concepts, shaping an idea into a plan and carrying an unfinished thought through to something usable. Why that matters →
It understands the task before I finish explaining it
This is the other part that has stood out to me. With many models — Gemini, or even GPT-6 Sol at its highest setting — I often still feel I have to explain the actual task twice. The first answer can be technically plausible, but it misses the real intention, a constraint from the project context or the reason a detail matters.
Claude has been the benchmark for this in my work. I started with Opus 4.6 and was struck by how well it understood what I meant, not merely the words I had typed. That ability kept improving through the current Opus 5, and Fable has it too: I can give it a rough voice transcription with sentence fragments, mistakes and missing context, and in perhaps 95 percent of cases it still identifies the right task and takes the right next steps.
Astra is the first model outside that Claude group that has reached the same level for me. It can work from an imperfect brief, infer the useful question and stay focused on the intended outcome without requiring a long second explanation. That is not a benchmark result. It is simply a practical difference you notice when the input comes from a real workday rather than a carefully written prompt.
The test was deliberately bounded
“Check the security of the whole application” is not a useful task for a model or a person. It produces a long, plausible document and little certainty.
The useful questions were concrete:
- Can a user cross a tenant boundary through this API flow?
- Do these RLS policies protect reads, writes and indirect access equally?
- Does this permission check still hold when a request arrives through a less obvious route?
- Which assumptions in this authentication flow need to be proven in code rather than trusted from the happy path?
That scope matters. It gives the review a target, makes the answer testable and leaves an audit trail when something needs to be checked manually.
Two independent reviews first, cross-review second
I do not show the first model’s answer to the second one at the start. That would make the second review less independent; a good-looking first explanation is exactly the kind of thing a model can accidentally anchor on.
Instead, Fable 5.1 and Astra receive the same brief and work through it separately. Both can inspect the relevant code, policies, API surface and project context. Both write their findings into the same result assignment — but neither begins from the other model’s conclusion.
Only then comes the second pass. Each model receives the other analysis and is asked to do something more useful than summarise it:
- confirm findings that are supported by the code;
- challenge findings that rest on an assumption;
- add a missing attack path or test case;
- name the points that still need a human decision or a manual verification.
The result is not one model sounding more confident than the other. It is a review with agreement, disagreement and evidence made visible.
What Astra changed in practice
The pattern I have seen is not that Astra produces a dramatically longer report. It is that it is more willing to stay with a narrow question until the conditions around it are clear.
In an authorization check, that can mean separating the rule itself from the assumptions it depends on: which identity reaches the policy, where that identity is established, whether an elevated route bypasses the expected check and whether the same restriction applies to writes as it does to reads. Those are exactly the gaps that are easy to miss when a review follows only the intended user flow.
Fable is still valuable in that process. It frequently establishes a strong first picture, and it is useful for turning findings into concrete remediation work. Astra has earned its place because it adds a genuinely independent perspective rather than a slightly different wording of the same answer.
Why the setup is practical in boostN
Running this manually would be tedious. You would have to set up separate prompts, preserve the first-pass results, distribute them to the other reviewers, keep track of which model challenged what and finally turn several conversations into a usable report.
boostN is built to make that setup repeatable.
I define the test once, choose the models and give each of them the same assignment. Their independent analyses land in one shared result card instead of disappearing into separate chats. From there, I can start a cross-review: each model checks the others’ findings against the original task and the available evidence.
That is the point of the workflow. It is not about collecting more model opinions. It is about making independent analysis and mutual review easy to run as one structured process.
The final result can then distinguish between confirmed findings, disputed findings, evidence, recommended fixes and the points that need a deeper human audit. That is far more useful than a single answer labelled “security review”.
The time and cost calculation
A proper check still takes time. The amount depends on the size of the system and on how much documentation the final report needs. In my current roughly five-hour usage window, I can usually work through two or three focused reviews of this kind. For a particularly tangled area, one end-to-end analysis may be the right limit.
There are not infinitely many high-risk areas in a typical product. RLS, API authorization, tenant boundaries, role changes, file access and a few important workflows cover a large part of the first pass. Even if I review only one of those areas a day, I can have the most important surfaces independently checked and cross-reviewed within a week.
For a 25-euro monthly plan, that is a striking amount of leverage. It is not comparable to commissioning a cybersecurity specialist. A specialist brings broader experience, deeper adversarial judgment and accountability that a model does not have. The model also cannot certify a system as secure.
But for a disciplined baseline check — and for deciding where specialist time is actually needed — this is already very good.
The boundary matters
I use this workflow to find and prioritise questions, not to declare a system safe. High-risk systems, a suspected incident or a compliance-critical release still need appropriate human security expertise.
My takeaway
I would not send Astra off to explore a production system without a defined task and clear boundaries. I would use it for exactly the opposite: a constrained review of a known surface, with a second model, a documented result and someone accountable for the final decision.
That is where it has been unexpectedly strong for me. Not as a replacement for security expertise, but as a practical way to make serious first-pass security work more thorough — and much easier to repeat.