Product

Deep Response Engine over AWS PrivateLink: Monitoring Your Infrastructure Without a Public Endpoint

How CloudThinker's Deep Response Engine reaches EKS, Prometheus, Redis, Postgres and Kafka inside a customer VPC over AWS PrivateLink, with no public endpoint. Covers AllowedPrincipals, read-only RBAC, one shared NLB, the AWS CLI setup and how to revoke access on your own.

WTWin Tran
privatelinkvpcendpointdeepresponseenginepulsesecuritynetworkingleastprivilegecompliancecloudthinker
Cover Image for Deep Response Engine over AWS PrivateLink: Monitoring Your Infrastructure Without a Public Endpoint

Deep Response Engine over AWS PrivateLink: Monitoring Your Infrastructure Without a Public Endpoint

How CloudThinker's Deep Response Engine reaches EKS, Prometheus, Redis and other production services inside a customer VPC without opening a single port to the internet.


The security review for a new monitoring vendor usually stalls on one line of the architecture diagram: an arrow from the vendor's cloud into your production VPC. The vendor needs Prometheus, the Kubernetes API, Redis and Postgres. Your security lead asks how that arrow gets in. The honest answer, for most tools, is a public load balancer, an IP allowlist and a long-lived token, and the review goes quiet.

This post explains how CloudThinker's Deep Response Engine (DRE) reaches the same systems with no public endpoint at all, what the customer controls at each hop, and how to set it up.

The public endpoint problem

Every third-party monitoring tool needs a way in. The default answer (expose an endpoint, allowlist an IP, hand over a token) is also the one your security team dreads signing off on.

A public endpoint is a permanent addition to your attack surface, whether or not anyone ever exploits it. It's a load balancer with a DNS name that resolves from anywhere, a target for credential stuffing and port scanners, and a line item your auditor will ask about by name. For teams pursuing SOC 2 or ISO 27001, "we expose Prometheus and our internal NLB to the public internet for a third-party AI vendor" isn't a sentence anyone wants in a control narrative. It invites exactly the finding compliance frameworks exist to prevent: an inbound path into production that isn't fully under your control.

The usual alternatives don't fit well either:

IP allowlists are only as good as the vendor's egress. A SaaS vendor's egress IPs are shared across its customers and change as it scales. The allowlist narrows who can reach the endpoint. It doesn't remove the endpoint.

VPC peering and VPNs grant too much. Peering connects whole CIDR ranges and needs non-overlapping address space with every vendor. A site-to-site VPN adds gateways to run and routes to keep correct. Both give the vendor a network path to far more than the five services it needs.

DRE (the incident investigation and remediation loop behind CloudThinker's AgenticOps platform) needs continuous access to real signal: Kubernetes events, Prometheus metrics, Redis health, Postgres performance and application latency from the API tier. That's a lot of surface to reach from outside a customer's VPC. The question isn't whether DRE needs connectivity. It's whether that connectivity has to cost you a public endpoint. It doesn't.


PrivateLink: the private backbone

CloudThinker runs outside every customer's VPC, as a multi-tenant SaaS platform. To let Pulse (DRE's signal ingestion layer) and the investigation agents reach a customer's infrastructure, we connect over AWS PrivateLink instead of the public internet.

01CloudThinker VPC

CloudThinkerCloudThinker

  • PulseSignal ingestion
  • Deep Response EngineTriage, investigate, remediate
  • Interface VPC EndpointAn ENI in CloudThinker’s VPC

Engineers approve from here

02The path

AWSPrivateLink

  • AWS private backboneNever leaves AWS
  • No public IPNothing to port-scan
  • No IGW, no NATThe internet isn’t in the path

One endpoint connection

03Customer VPC

Your account

  • AllowedPrincipalsOnly the CloudThinker IAM role
  • VPC Endpoint ServiceYou own it, you can revoke it
  • Internal NLBShared, per-port target groups

Your off switch, not ours

Every hop between CloudThinker and a customer VPC stays on AWS’s private backbone, and the customer’s endpoint service decides who may connect at all.

The mechanics matter more than the term. CloudThinker maintains an Interface VPC Endpoint with an ENI inside CloudThinker's own VPC. Traffic from that endpoint travels over AWS's private backbone to a VPC Endpoint Service the customer creates and owns. At no point does a packet get a public IP address, traverse an internet gateway or leave AWS's network. There's no NAT gateway in the path and no firewall to route around. The internet simply isn't part of the path.

That single property collapses most of the attack surface argument. You can't port-scan a service that has no public IP. You can't credential-stuff an endpoint that only resolves inside AWS's backbone. And for compliance, "traffic never traverses the public internet" is a control statement your auditor can verify directly against the VPC Endpoint Service configuration, not a claim resting on firewall rules staying correct forever.

The connection is also one-directional. PrivateLink lets the consumer (CloudThinker's endpoint) open connections to the services behind your NLB. It doesn't give anything in your VPC a route back into CloudThinker's network, and it doesn't expose anything in your VPC that isn't behind that NLB.

The customer keeps the steering wheel throughout. AllowedPrincipals on the VPC Endpoint Service is the authorization gate: only the specific IAM role CloudThinker connects with is on that list. No other AWS account and no other role can even request a connection. The customer accepts the endpoint connection, and can revoke it at any time by editing that list or deleting the endpoint service outright. This is infrastructure the customer provisions and can tear down unilaterally. CloudThinker never holds the only copy of the off switch.


Least privilege, even on a private path

A private network path answers "can anything on the internet reach this?" It doesn't answer "what can CloudThinker do once connected?" That's a separate question, and we treat it as one.

01Network

Who can connect

PrivateLink plus AllowedPrincipals: one IAM role, no public path, revocable by you.

02Credentials

What a connection can do

get, list and watch only. A misused credential can read metrics, not delete a deployment.

Network isolation and credential scoping are separate controls. DRE relies on both.

Every credential DRE uses to reach customer infrastructure is scoped down before it's ever used. The Kubernetes ClusterRole granted to CloudThinker's service account allows get, list and watch (full read access to observe cluster state, events and logs) and nothing else. No create, no update, no delete, no patch. A minimal version looks like this:

apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: cloudthinker-readonly
rules:
  - apiGroups: ['', 'apps', 'batch', 'autoscaling', 'networking.k8s.io']
    resources: ['*']
    verbs: ['get', 'list', 'watch']
  - apiGroups: ['metrics.k8s.io']
    resources: ['pods', 'nodes']
    verbs: ['get', 'list']

If your policy keeps Secrets out of reach of any vendor, replace the wildcard with an explicit resource list that leaves secrets out. The IAM role authorized on the VPC Endpoint Service carries the same shape: read-scoped policies, no wildcard actions, no administrative grants.

This matters even inside a private network because PrivateLink controls who can connect, not what a connection is allowed to do. If DRE's credentials were ever misused, replayed or targeted, a read-only role limits the blast radius to "an attacker can see your metrics", not "an attacker can delete your deployment". Network isolation and credential scoping are independent controls, and DRE relies on both rather than treating either as enough on its own.


One shared NLB, not one endpoint per service

DRE needs signal from more than one system per cluster: typically the API tier, the primary database, the cache layer, the message queue, and Prometheus for metrics. A dedicated Network Load Balancer per service, each behind its own endpoint, is both an operational and a cost problem: more NLBs to provision, patch and monitor, and more hourly charges for infrastructure that's mostly idle.

01One endpoint

AWSinternal-nlb

  • Routes by portOne listener per service
  • One target group per namespaceIP targets on the cluster

One artifact to approve

02EKS: customer-production

Five services

  • ns: apiapi-svc
  • PostgreSQLns: databasepostgres-svc
  • Redisns: cacheredis-svc
  • Kafkans: messagingkafka-svc
  • Prometheusns: monitoringprometheus-svc

Fewer endpoints to pay for and audit

One shared NLB in front of the endpoint service instead of one load balancer per service.

In practice, one internal NLB in front of the VPC Endpoint Service is enough. It routes by port to a separate target group per namespace (api, database, cache, messaging, monitoring), so five services share one load balancer and one endpoint connection instead of five. From the customer's side, that's a single artifact to review, approve and audit, rather than five separate attack surfaces to track.


Setting it up

01Front

Internal NLB

Listeners per port, target groups per namespace.

02Expose

Endpoint service

Created on the NLB, acceptance required.

03Authorize

One principal

Add only the CloudThinker IAM role to AllowedPrincipals.

04Accept

The connection

You approve the endpoint request. Revoke it any time.

The 30 to 60 minutes with your network team, in four steps.

The Connections Guide has the full walkthrough. The shape of it, with the AWS CLI, looks like this. Values in angle brackets are yours, and CloudThinker provides the IAM role ARN to authorize.

1. Create the endpoint service on your internal NLB, with acceptance required:

aws ec2 create-vpc-endpoint-service-configuration \
  --network-load-balancer-arns <internal-nlb-arn> \
  --acceptance-required

2. Authorize exactly one principal:

aws ec2 modify-vpc-endpoint-service-permissions \
  --service-id <vpce-svc-id> \
  --add-allowed-principals <cloudthinker-role-arn>

3. Accept the endpoint connection once CloudThinker requests it:

aws ec2 describe-vpc-endpoint-connections \
  --filters Name=service-id,Values=<vpce-svc-id>

aws ec2 accept-vpc-endpoint-connections \
  --service-id <vpce-svc-id> \
  --vpc-endpoint-ids <vpce-id>

Revoking is the same tools in reverse. Remove the principal with --remove-allowed-principals, reject the connection with aws ec2 reject-vpc-endpoint-connections, or delete the service with aws ec2 delete-vpc-endpoint-service-configurations. Any one of them cuts DRE off, and none of them needs CloudThinker's cooperation.


What DRE still does with less access

None of this trades away investigative capability for safety. Pulse still ingests the full signal stream from every connected service, DRE still clusters noisy alerts into coherent incidents, and when a pattern crosses the escalation threshold, the investigation loop still runs full root-cause analysis across the same telemetry.

Here's what that looks like on a real class of incident. Checkout latency climbs at 14:20, and DRE has three readings to test, all through read-only calls over the private link:

  • The cache is degraded. If true, Redis INFO shows evictions or a jump in instantaneous_ops_per_sec and latency on the cache service. DRE reads Redis health on the cache port and finds eviction counts flat. Ruled out.
  • The database is saturated. If true, Postgres pg_stat_activity shows connections near max_connections and long-running queries. Connections are at 40% with no long waits. Ruled out.
  • A change landed. If true, the Kubernetes API shows a new ReplicaSet for the API tier near 14:20. kubectl get rs -n api (a list, which the ClusterRole allows) shows a rollout at 14:17, and Prometheus shows the p95 step change three minutes later.

The finding names the rollout as the likely cause and proposes a rollback. Nothing in that investigation needed a write permission or a public endpoint.

Remediation follows the Auto Mode level the customer configures: Manual, where DRE proposes a fix and an engineer approves it, or Auto, where DRE executes runbook-defined actions directly and escalates anything outside them. Either way, DRE's own connection into the customer's environment stays read-only for observability. Write actions during remediation are performed through the customer's own scoped automation, not by widening what DRE's monitoring credentials can touch. The engineer stays in the loop at the level they choose, reviewing and approving from outside the VPC, never needing standing access inside it.


Why this beats the usual options

Public endpoint VPC peering or VPN PrivateLink with DRE
Public IP to defend Yes No No
What the vendor can reach Whatever is behind the endpoint Routed CIDR ranges Only the services behind one NLB
Who can connect Anyone who passes the allowlist and token Anything on the peered network One IAM role on AllowedPrincipals
Address planning None Non-overlapping CIDRs required None
Revocation Rotate tokens, edit firewall rules Tear down peering or tunnels Remove the principal or reject the connection
Audit statement "Firewall rules restrict access" "Routes and NACLs restrict access" "Traffic never leaves AWS's backbone"

The deeper difference is where the control sits. With a public endpoint, safety depends on rules staying correct forever. With PrivateLink, the default is no path at all, and the only path that exists is one you created, scoped to one role and can delete on your own.


Getting connected

Public endpoints trade a little setup friction for a permanent increase in attack surface. PrivateLink asks for 30 to 60 minutes with your network team once, in exchange for a connection that never has a public IP to defend, and an access grant your security team can revoke on their own timeline.

For DRE, that means full signal visibility and full investigative capability into EKS, Prometheus, Redis, Postgres and Kafka, without opening a single inbound rule to the internet.

Once connected, you can check the setup in plain language:

"List every service you can reach in our production account over PrivateLink, and the permissions you hold on each."

"Checkout p95 has been climbing since 14:20. Investigate across Redis, Postgres and recent rollouts, and propose a fix. Manual mode, change nothing."

→ Read the Connections Guide for VPC Endpoint setup, step by step

→ Contact Us for PrivateLink Setup for enterprise connectivity with dedicated support

Related reading