Deep Response Engine over AWS PrivateLink: Monitoring Your Infrastructure Without a Public Endpoint
How CloudThinker's Deep Response Engine reaches EKS, Prometheus, Redis and other production services inside a customer VPC without opening a single port to the internet.
The security review for a new monitoring vendor usually stalls on one line of the architecture diagram: an arrow from the vendor's cloud into your production VPC. The vendor needs Prometheus, the Kubernetes API, Redis and Postgres. Your security lead asks how that arrow gets in. The honest answer, for most tools, is a public load balancer, an IP allowlist and a long-lived token, and the review goes quiet.
This post explains how CloudThinker's Deep Response Engine (DRE) reaches the same systems with no public endpoint at all, what the customer controls at each hop, and how to set it up.
The public endpoint problem
Every third-party monitoring tool needs a way in. The default answer (expose an endpoint, allowlist an IP, hand over a token) is also the one your security team dreads signing off on.
A public endpoint is a permanent addition to your attack surface, whether or not anyone ever exploits it. It's a load balancer with a DNS name that resolves from anywhere, a target for credential stuffing and port scanners, and a line item your auditor will ask about by name. For teams pursuing SOC 2 or ISO 27001, "we expose Prometheus and our internal NLB to the public internet for a third-party AI vendor" isn't a sentence anyone wants in a control narrative. It invites exactly the finding compliance frameworks exist to prevent: an inbound path into production that isn't fully under your control.
The usual alternatives don't fit well either:
IP allowlists are only as good as the vendor's egress. A SaaS vendor's egress IPs are shared across its customers and change as it scales. The allowlist narrows who can reach the endpoint. It doesn't remove the endpoint.
VPC peering and VPNs grant too much. Peering connects whole CIDR ranges and needs non-overlapping address space with every vendor. A site-to-site VPN adds gateways to run and routes to keep correct. Both give the vendor a network path to far more than the five services it needs.
DRE (the incident investigation and remediation loop behind CloudThinker's AgenticOps platform) needs continuous access to real signal: Kubernetes events, Prometheus metrics, Redis health, Postgres performance and application latency from the API tier. That's a lot of surface to reach from outside a customer's VPC. The question isn't whether DRE needs connectivity. It's whether that connectivity has to cost you a public endpoint. It doesn't.
PrivateLink: the private backbone
CloudThinker runs outside every customer's VPC, as a multi-tenant SaaS platform. To let Pulse (DRE's signal ingestion layer) and the investigation agents reach a customer's infrastructure, we connect over AWS PrivateLink instead of the public internet.
01CloudThinker VPC
CloudThinker
- PulseSignal ingestion
- Deep Response EngineTriage, investigate, remediate
- Interface VPC EndpointAn ENI in CloudThinker’s VPC
Engineers approve from here
02The path
PrivateLink
- AWS private backboneNever leaves AWS
- No public IPNothing to port-scan
- No IGW, no NATThe internet isn’t in the path
One endpoint connection
03Customer VPC
Your account
- AllowedPrincipalsOnly the CloudThinker IAM role
- VPC Endpoint ServiceYou own it, you can revoke it
- Internal NLBShared, per-port target groups
Your off switch, not ours
The mechanics matter more than the term. CloudThinker maintains an Interface VPC Endpoint with an ENI inside CloudThinker's own VPC. Traffic from that endpoint travels over AWS's private backbone to a VPC Endpoint Service the customer creates and owns. At no point does a packet get a public IP address, traverse an internet gateway or leave AWS's network. There's no NAT gateway in the path and no firewall to route around. The internet simply isn't part of the path.
That single property collapses most of the attack surface argument. You can't port-scan a service that has no public IP. You can't credential-stuff an endpoint that only resolves inside AWS's backbone. And for compliance, "traffic never traverses the public internet" is a control statement your auditor can verify directly against the VPC Endpoint Service configuration, not a claim resting on firewall rules staying correct forever.
The connection is also one-directional. PrivateLink lets the consumer (CloudThinker's endpoint) open connections to the services behind your NLB. It doesn't give anything in your VPC a route back into CloudThinker's network, and it doesn't expose anything in your VPC that isn't behind that NLB.
The customer keeps the steering wheel throughout. AllowedPrincipals on the VPC Endpoint Service is the authorization gate: only the specific IAM role CloudThinker connects with is on that list. No other AWS account and no other role can even request a connection. The customer accepts the endpoint connection, and can revoke it at any time by editing that list or deleting the endpoint service outright. This is infrastructure the customer provisions and can tear down unilaterally. CloudThinker never holds the only copy of the off switch.
Least privilege, even on a private path
A private network path answers "can anything on the internet reach this?" It doesn't answer "what can CloudThinker do once connected?" That's a separate question, and we treat it as one.
01Network
Who can connect
PrivateLink plus AllowedPrincipals: one IAM role, no public path, revocable by you.
02Credentials
What a connection can do
get, list and watch only. A misused credential can read metrics, not delete a deployment.
Every credential DRE uses to reach customer infrastructure is scoped down before it's ever used. The Kubernetes ClusterRole granted to CloudThinker's service account allows get, list and watch (full read access to observe cluster state, events and logs) and nothing else. No create, no update, no delete, no patch. A minimal version looks like this:
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: cloudthinker-readonly
rules:
- apiGroups: ['', 'apps', 'batch', 'autoscaling', 'networking.k8s.io']
resources: ['*']
verbs: ['get', 'list', 'watch']
- apiGroups: ['metrics.k8s.io']
resources: ['pods', 'nodes']
verbs: ['get', 'list']
If your policy keeps Secrets out of reach of any vendor, replace the wildcard with an explicit resource list that leaves secrets out. The IAM role authorized on the VPC Endpoint Service carries the same shape: read-scoped policies, no wildcard actions, no administrative grants.
This matters even inside a private network because PrivateLink controls who can connect, not what a connection is allowed to do. If DRE's credentials were ever misused, replayed or targeted, a read-only role limits the blast radius to "an attacker can see your metrics", not "an attacker can delete your deployment". Network isolation and credential scoping are independent controls, and DRE relies on both rather than treating either as enough on its own.
One shared NLB, not one endpoint per service
DRE needs signal from more than one system per cluster: typically the API tier, the primary database, the cache layer, the message queue, and Prometheus for metrics. A dedicated Network Load Balancer per service, each behind its own endpoint, is both an operational and a cost problem: more NLBs to provision, patch and monitor, and more hourly charges for infrastructure that's mostly idle.
01One endpoint
internal-nlb
- Routes by portOne listener per service
- One target group per namespaceIP targets on the cluster
One artifact to approve
02EKS: customer-production
Five services
- ns: apiapi-svc
ns: databasepostgres-svc
ns: cacheredis-svc
ns: messagingkafka-svc
ns: monitoringprometheus-svc
Fewer endpoints to pay for and audit
In practice, one internal NLB in front of the VPC Endpoint Service is enough. It routes by port to a separate target group per namespace (api, database, cache, messaging, monitoring), so five services share one load balancer and one endpoint connection instead of five. From the customer's side, that's a single artifact to review, approve and audit, rather than five separate attack surfaces to track.
Setting it up
01Front
Internal NLB
Listeners per port, target groups per namespace.
02Expose
Endpoint service
Created on the NLB, acceptance required.
03Authorize
One principal
Add only the CloudThinker IAM role to AllowedPrincipals.
04Accept
The connection
You approve the endpoint request. Revoke it any time.
The Connections Guide has the full walkthrough. The shape of it, with the AWS CLI, looks like this. Values in angle brackets are yours, and CloudThinker provides the IAM role ARN to authorize.
1. Create the endpoint service on your internal NLB, with acceptance required:
aws ec2 create-vpc-endpoint-service-configuration \
--network-load-balancer-arns <internal-nlb-arn> \
--acceptance-required
2. Authorize exactly one principal:
aws ec2 modify-vpc-endpoint-service-permissions \
--service-id <vpce-svc-id> \
--add-allowed-principals <cloudthinker-role-arn>
3. Accept the endpoint connection once CloudThinker requests it:
aws ec2 describe-vpc-endpoint-connections \
--filters Name=service-id,Values=<vpce-svc-id>
aws ec2 accept-vpc-endpoint-connections \
--service-id <vpce-svc-id> \
--vpc-endpoint-ids <vpce-id>
Revoking is the same tools in reverse. Remove the principal with --remove-allowed-principals, reject the connection with aws ec2 reject-vpc-endpoint-connections, or delete the service with aws ec2 delete-vpc-endpoint-service-configurations. Any one of them cuts DRE off, and none of them needs CloudThinker's cooperation.
What DRE still does with less access
None of this trades away investigative capability for safety. Pulse still ingests the full signal stream from every connected service, DRE still clusters noisy alerts into coherent incidents, and when a pattern crosses the escalation threshold, the investigation loop still runs full root-cause analysis across the same telemetry.
Here's what that looks like on a real class of incident. Checkout latency climbs at 14:20, and DRE has three readings to test, all through read-only calls over the private link:
- The cache is degraded. If true, Redis
INFOshows evictions or a jump ininstantaneous_ops_per_secand latency on the cache service. DRE reads Redis health on thecacheport and finds eviction counts flat. Ruled out. - The database is saturated. If true, Postgres
pg_stat_activityshows connections nearmax_connectionsand long-running queries. Connections are at 40% with no long waits. Ruled out. - A change landed. If true, the Kubernetes API shows a new ReplicaSet for the API tier near 14:20.
kubectl get rs -n api(alist, which the ClusterRole allows) shows a rollout at 14:17, and Prometheus shows the p95 step change three minutes later.
The finding names the rollout as the likely cause and proposes a rollback. Nothing in that investigation needed a write permission or a public endpoint.
Remediation follows the Auto Mode level the customer configures: Manual, where DRE proposes a fix and an engineer approves it, or Auto, where DRE executes runbook-defined actions directly and escalates anything outside them. Either way, DRE's own connection into the customer's environment stays read-only for observability. Write actions during remediation are performed through the customer's own scoped automation, not by widening what DRE's monitoring credentials can touch. The engineer stays in the loop at the level they choose, reviewing and approving from outside the VPC, never needing standing access inside it.
Why this beats the usual options
| Public endpoint | VPC peering or VPN | PrivateLink with DRE | |
|---|---|---|---|
| Public IP to defend | Yes | No | No |
| What the vendor can reach | Whatever is behind the endpoint | Routed CIDR ranges | Only the services behind one NLB |
| Who can connect | Anyone who passes the allowlist and token | Anything on the peered network | One IAM role on AllowedPrincipals |
| Address planning | None | Non-overlapping CIDRs required | None |
| Revocation | Rotate tokens, edit firewall rules | Tear down peering or tunnels | Remove the principal or reject the connection |
| Audit statement | "Firewall rules restrict access" | "Routes and NACLs restrict access" | "Traffic never leaves AWS's backbone" |
The deeper difference is where the control sits. With a public endpoint, safety depends on rules staying correct forever. With PrivateLink, the default is no path at all, and the only path that exists is one you created, scoped to one role and can delete on your own.
Getting connected
Public endpoints trade a little setup friction for a permanent increase in attack surface. PrivateLink asks for 30 to 60 minutes with your network team once, in exchange for a connection that never has a public IP to defend, and an access grant your security team can revoke on their own timeline.
For DRE, that means full signal visibility and full investigative capability into EKS, Prometheus, Redis, Postgres and Kafka, without opening a single inbound rule to the internet.
Once connected, you can check the setup in plain language:
"List every service you can reach in our production account over PrivateLink, and the permissions you hold on each."
"Checkout p95 has been climbing since 14:20. Investigate across Redis, Postgres and recent rollouts, and propose a fix. Manual mode, change nothing."
→ Read the Connections Guide for VPC Endpoint setup, step by step
→ Contact Us for PrivateLink Setup for enterprise connectivity with dedicated support
