The four buckets every AWS SOC 2 team ends up in
If you passed SOC 2 Type 2 and don't have a dedicated DevSecOps hire, your open findings almost certainly cluster into the same four buckets. Here's what they look like, and why the finicky ones quietly break renewal.
The four buckets every AWS SOC 2 team ends up in
By David Thompson, CTO & co-founder, Tudovu
Every AWS team I've watched pass SOC 2 Type 2 ends up with the same open findings a few months later. I've run the compliance, security, and infrastructure for two different successful exits. Different stacks, different auditors, and the post-audit findings list has come back the same shape every time.
If you're on AWS, passed Type 2 in the last 18 months, and don't have a dedicated DevSecOps hire, I can probably describe your open findings without seeing your account. Companies of the same size drift the same way.
Below is that pattern: four buckets, the finding everyone misses in each, the one that fights back when you fix it, and why the tools that find all of it stop short of fixing it. The short version: detection is solved. The gap is getting fixes into your infrastructure code without fighting your own pipeline.
The shape of the surface
Across the teams I've worked with, open findings sort into four buckets in roughly the same proportions. Logging, monitoring, and alerting is always the biggest. IAM, network, and data protection run roughly even behind it.
A Type 2 report covers a window that has already closed. Once it lands, new services keep shipping from the same templates, someone grants prod access at 2 a.m., and nobody owns posture anymore. The drift sits there until renewal prep, which is the worst time to find it.
Bucket 1: IAM drift
Maps to: SOC 2 CC6.1, CC6.3
What the scanner shows you:
- Admin-adjacent AWS-managed policies attached to users who don't need them
- Inline policies that escalate through iam:PassRole plus broad service actions
- Access keys past 90-day rotation
- Users with directly attached policies instead of group-based access
- Root account with recent activity
The one everyone misses is an inline policy that pairs iam:PassRole on * with broad Lambda permissions. Scanners rate it medium. In practice it's enough to create a function, hand it a more privileged role, and invoke it to mint access keys for an admin user. That's the least-privilege failure CC6.3 exists to catch.
The finicky one is directly attached policies. Each looks like a two-hour cleanup PR. Half of them exist because someone needed prod access at 2 a.m. six months ago, and each has a person and a reason behind it. Removing them turns into a week of change management, because every removal is a conversation.
Bucket 2: Network exposure
Maps to: SOC 2 CC6.6
What the scanner shows you:
- Subnets auto-assigning public IPs on launch
- Network ACLs allowing internet ingress on 22 or 3389
- Security groups with 0.0.0.0/0 open on ports that shouldn't be
- Default VPC security groups never locked down
- EC2 instances with public IPs they don't need
The one everyone misses is an ECS service whose tasks get public IPs because the template it was copied from set AssignPublicIp to ENABLED. CloudFormation defaults it to disabled, but plenty of starter templates flip it on so tasks in a public subnet can pull images without a NAT gateway. Nobody on your team decided to make the service public. It happens again every time a new service ships from that template.
The finicky one is the default VPC security group. Every scanner flags it, and you can't delete it. The fix is stripping its rules, which breaks anything quietly relying on it: an instance launched without an explicit group, or a database someone attached to "default" years ago. Inventory what references it before you touch it. Most teams file an exception and hope the next auditor doesn't push. Some auditors do.
Bucket 3: Logging, monitoring, and alerting
Maps to: SOC 2 CC7.2, CC7.3
What the scanner shows you:
- CloudWatch log retention shorter than 365 days
- Log groups not encrypted with a customer-managed KMS key
- ELBv2 access logging off
- RDS logs not exported to CloudWatch
- WAFv2 logging off
- EventBridge → SNS alerting missing for EC2 stop/terminate, AWS Health, Auto Scaling failures, and RDS events
This is the biggest bucket, and it's where "just enable it" goes wrong most often. Retention and KMS get the attention, while the new log sources quietly run up the bill. CloudWatch bills about $0.50 per GB on the way in (us-east-1) and about $0.03 per GB-month to store it, so the real cost is every log source you switch on to clear a finding. Send high-volume sources to S3 and keep CloudWatch for what you query and alert on. Every compliance checkbox in this bucket has an invoice attached to it.
The one everyone misses is the log groups Lambda creates for itself. The first time a function runs, Lambda creates its log group with no retention and no customer-managed key. Every new function brings the finding back unless your module declares that log group first.
The finicky one is EventBridge → SNS wiring for six event sources. Each needs its own rule and event pattern, a topic policy that lets EventBridge publish, and a confirmed subscription on the other end. If the topic uses a customer-managed KMS key, which the same scanner wants, EventBridge can't publish until the key policy allows it. Nothing tells you the alerts aren't arriving unless someone watches the rule's failed invocations. That's a sprint of work, and in the audit report it shows up as six red checks turning green.
Bucket 4: Data protection and recovery
Maps to: SOC 2 CC6.7, CC7.5, and A1.2 if Availability is in your scope
What the scanner shows you:
- S3 account-level Block Public Access not enforced
- Bucket-level Block Public Access partial (2 of the 4 settings on)
- S3 server access logging off
- RDS Multi-AZ off on production databases
- RDS deletion protection off
- Secrets Manager rotation off on sampled secrets
The one everyone misses is what rotation does to live Lambdas that cache a secret. With single-user rotation, the password changes under them and every cached connection starts failing. Switch those secrets to the alternating-users strategy, or make the functions refetch on an auth failure, before you turn rotation on. If you've done it before, it's a two-hour fix. If you haven't, it's an outage.
The finicky one is Block Public Access. The scanner checks it at the account level and the bucket level, and account-level wins: once it's on, no bucket in the account can be public, whatever its own settings say. That breaks anything legitimately public, like marketing assets, hosted docs, or a static site served straight from S3. Inventory first, move public content behind CloudFront with origin access control, then enforce at the account level.
What SOC 2 actually requires
Everything above is what a SOC 2-mapped check set flags, and the check set is stricter than the standard. SOC 2 criteria describe outcomes, like restricting logical access and detecting security events. They don't mention 365 days, Multi-AZ, or customer-managed keys. Those numbers come from benchmarks like CIS, your own policies, and what your auditor expects.
That changes how you triage. Some findings get fixed. Some get a documented risk acceptance. Some fall out of scope entirely, like the A1 items when Availability isn't in your report, since Security is the only required category. Sort the queue that way before anyone writes code.
Why the tooling stops short of the fix
Detection is a solved problem. Vanta, Drata, Secureframe, Wiz, Prowler, Security Hub: pick any of them and they'll find all of this. GRC platforms handle the automated checks well. They stall on anything that needs a human to do the work, and every finding above needs an engineer to write a change, get it reviewed, and ship it.
"Auto-remediation" doesn't close that queue either, at least in any shop that runs infrastructure as code. A fix flipped in the console is drift the moment it lands. The next terraform plan sees it and tries to revert it, or someone ends up doing state surgery to reconcile it. A tool that flips settings in the console is fighting your pipeline. A tool that changes the source your pipeline deploys from is working with it.
What actually closes the loop
The teams I've seen keep their posture stable after Type 2 converge on the same workflow, whatever tooling they use to get there.
Findings land as proposed changes to the infrastructure code, in the repo the team already works in. An engineer reads the diff, requests changes if something's off, and merges. The finding closes because the deployed configuration changed. The commit, the review, and the merge timestamp become evidence for the next audit, with no screenshot folder to assemble.
Access changes go through the same loop, so the 2 a.m. grant arrives as a pull request with a name and an expiry date attached. Nothing gets flipped in the console behind the pipeline's back. Nothing ships without an engineer approving it. Whichever IaC layer owns your infrastructure keeps owning it.
The math
Without a DevSecOps hire, working through a queue like this usually takes two engineers about six weeks. At roughly $200K a year loaded, that's around $45K in salary, plus six weeks of roadmap those engineers didn't ship. That can easily exceed the audit itself: an independent directory of 192 audit firms puts a Type 2 with a specialist CPA firm at [$15,500 to $50,000](https://soc2auditors.org/soc-2-audit-cost/).
A dedicated hire runs about $200K a year loaded. That buys an owner for the queue, but the fire drill still repeats every cycle, because the templates and defaults that filled the buckets haven't changed.
The third option is fixing the source: correct the templates, route every fix and every access change through the same review your product code gets, and let the merge history serve as evidence. The buckets still refill, slowly enough that nobody has to drop the roadmap to empty them.
David Thompson is CTO and co-founder of Tudovu. He has run compliance, security, and infrastructure through two successful exits. If your team is staring at a finding queue between now and Type 2 renewal, [book a Compliance Engineering Walkthrough with our team](https://calendly.com/easmond-tsewole) and we'll work through a real finding together on the call.