Skip to content

GCO Global API Gateway — API spec sheet

API Gateway api-gateway-global · OpenAPI 3.0.3 · API version 1.0.0 · 14 endpoints · 0 schemas

Authenticated global aggregation API for GCO

The single authenticated entry point for GCO. Every request is signed with AWS SigV4 and authorized by IAM; the Lambda proxies behind the routes add the request-bound HMAC envelope that the in-cluster services require, so this API is the only supported caller of the cluster besides the regional bridges. /api/v1/{proxy+} and /inference/{proxy+} reach the regions over Global Accelerator; /api/v1/global/* fans out to every regional API Gateway; /studio/* exists only with the analytics environment.

Servers

  • https://{api_id}.execute-api.{region}.{url_suffix}/{stage}

Stage prod of the deployed REST API. The full URL is the ApiEndpoint output of the gco-api-gateway stack.

Variable Default Description
api_id "<api-id>" REST API id assigned at deploy time.
region "us-east-2" deployment_regions.api_gateway in cdk.json.
stage "prod" —
url_suffix "amazonaws.com" The partition's DNS suffix (${AWS::URLSuffix}).

Deployment

  • Endpoint type: EDGE
  • Stage: prod — throttling 1000 req/s (burst 2000), execution logging INFO, data trace off, metrics on, X-Ray tracing on, access logs on
  • Conditional routes: operations marked Only when exist under:
  • analytics: Present only when analytics_environment.enabled is true in cdk.json.
  • global-accelerator: Present only in the commercial aws partition, where Global Accelerator provides the data path to the regional ALBs. Elsewhere the global API keeps only its aggregate routes and workload traffic uses each region's bridge.

Resource policy

Statements deployed on the REST API (${AWS::AccountId} is the deploying account):

[
  {
    "Action": "execute-api:Invoke",
    "Condition": {
      "StringEquals": {
        "aws:PrincipalAccount": "${AWS::AccountId}"
      }
    },
    "Effect": "Allow",
    "Principal": {
      "AWS": "*"
    },
    "Resource": "execute-api:/*"
  },
  {
    "Action": "execute-api:Invoke",
    "Effect": "Allow",
    "Principal": {
      "AWS": "*"
    },
    "Resource": "execute-api:/*/GET/studio/*"
  }
]

WAF

Priority Rule Action Managed rule group
0 PerIPRateLimit block —
1 NonInferenceBodySizeLimit block —
2 AWSManagedRulesCommonRuleSet group default AWSManagedRulesCommonRuleSet
3 AWSManagedRulesKnownBadInputsRuleSet group default AWSManagedRulesKnownBadInputsRuleSet
4 AWSManagedRulesSQLiRuleSet group default AWSManagedRulesSQLiRuleSet
5 AWSManagedRulesLinuxRuleSet group default AWSManagedRulesLinuxRuleSet
6 AWSManagedRulesAmazonIpReputationList group default AWSManagedRulesAmazonIpReputationList
7 AWSManagedRulesAnonymousIpList group default AWSManagedRulesAnonymousIpList

Security schemes

cognito

  • Type: apiKey — header Authorization (cognito_user_pools)

ID token issued by the Studio Cognito user pool, validated by the <project>-studio-cognito-authorizer authorizer.

sigv4

  • Type: apiKey — header Authorization (awsSigv4)

AWS Signature Version 4 for service execute-api, signed with IAM credentials of the deploying account that carry execute-api:Invoke on this API (for example awscurl --service execute-api --region <region>).

Endpoints

Method Path Summary Tags
GET /api/v1/global/health Health of every region's cluster global
GET /api/v1/global/jobs List jobs across every region global
DELETE /api/v1/global/jobs Bulk-delete jobs across every region global
GET /api/v1/global/status Aggregated status of every region global
GET /api/v1/{proxy+} Read a control-plane resource control-plane
POST /api/v1/{proxy+} Create a control-plane resource or submit a job control-plane
PUT /api/v1/{proxy+} Replace a control-plane resource control-plane
PATCH /api/v1/{proxy+} Update a control-plane resource control-plane
DELETE /api/v1/{proxy+} Delete a control-plane resource control-plane
GET /inference/{proxy+} Read an inference endpoint's health, info or model list inference
POST /inference/{proxy+} Generate: forward a completion request and stream the response inference
HEAD /inference/{proxy+} Probe an inference endpoint inference
GET /studio/callback OAuth redirect landing page (stub) studio
GET /studio/login Presigned SageMaker Studio login URL studio

Endpoint details

GET /api/v1/global/health

Fans out to /api/v1/health on every regional API Gateway and returns one health object per region.

  • Operation ID: getApiV1GlobalHealth
  • Tags: global
  • Security: sigv4
  • Integration: aws_proxy, 29 s integration timeout

Backend

  • Lambda: cross-region-aggregator (lambda/cross-region-aggregator/handler.py; python3.14, handler.lambda_handler, 29 s timeout)
  • Behaviour: Discovers every -regional-api- stack's RegionalApiEndpoint output, fans the request out to each regional API with SigV4 and merges the answers into one response.
  • Then:
    1. Each region's API Gateway bridge, called with SigV4 by the aggregator's own execution role (GET/DELETE api/v1/jobs, GET api/v1/health, api/v1/status, api/v1/policy) — api-gateway-regional
    2. regional-api-proxy VPC Lambda: resolves the ALB from SSM //alb-hostname- and signs the HMAC envelope
    3. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    4. manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) — manifest-processor, health-monitor

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
500 Internal error in the proxy Lambda or the backend. —

GET /api/v1/global/jobs

Fans out to /api/v1/jobs on every regional API Gateway and merges the results. Requires the regional bridges to be deployed; each region's answer is attributed to it in the response.

  • Operation ID: getApiV1GlobalJobs
  • Tags: global
  • Security: sigv4
  • Integration: aws_proxy, 29 s integration timeout

Backend

  • Lambda: cross-region-aggregator (lambda/cross-region-aggregator/handler.py; python3.14, handler.lambda_handler, 29 s timeout)
  • Behaviour: Discovers every -regional-api- stack's RegionalApiEndpoint output, fans the request out to each regional API with SigV4 and merges the answers into one response.
  • Then:
    1. Each region's API Gateway bridge, called with SigV4 by the aggregator's own execution role (GET/DELETE api/v1/jobs, GET api/v1/health, api/v1/status, api/v1/policy) — api-gateway-regional
    2. regional-api-proxy VPC Lambda: resolves the ALB from SSM //alb-hostname- and signs the HMAC envelope
    3. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    4. manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) — manifest-processor, health-monitor

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
500 Internal error in the proxy Lambda or the backend. —

DELETE /api/v1/global/jobs

Fans out to /api/v1/jobs on every regional API Gateway and merges the results. Requires the regional bridges to be deployed; each region's answer is attributed to it in the response.

  • Operation ID: deleteApiV1GlobalJobs
  • Tags: global
  • Security: sigv4
  • Integration: aws_proxy, 29 s integration timeout

Backend

  • Lambda: cross-region-aggregator (lambda/cross-region-aggregator/handler.py; python3.14, handler.lambda_handler, 29 s timeout)
  • Behaviour: Discovers every -regional-api- stack's RegionalApiEndpoint output, fans the request out to each regional API with SigV4 and merges the answers into one response.
  • Then:
    1. Each region's API Gateway bridge, called with SigV4 by the aggregator's own execution role (GET/DELETE api/v1/jobs, GET api/v1/health, api/v1/status, api/v1/policy) — api-gateway-regional
    2. regional-api-proxy VPC Lambda: resolves the ALB from SSM //alb-hostname- and signs the HMAC envelope
    3. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    4. manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) — manifest-processor, health-monitor

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
500 Internal error in the proxy Lambda or the backend. —

GET /api/v1/global/status

Fans out to /api/v1/status on every regional API Gateway (answered by the manifest-processor) and merges templates, webhooks, resource limits and allowed namespaces per region.

  • Operation ID: getApiV1GlobalStatus
  • Tags: global
  • Security: sigv4
  • Integration: aws_proxy, 29 s integration timeout

Backend

  • Lambda: cross-region-aggregator (lambda/cross-region-aggregator/handler.py; python3.14, handler.lambda_handler, 29 s timeout)
  • Behaviour: Discovers every -regional-api- stack's RegionalApiEndpoint output, fans the request out to each regional API with SigV4 and merges the answers into one response.
  • Then:
    1. Each region's API Gateway bridge, called with SigV4 by the aggregator's own execution role (GET/DELETE api/v1/jobs, GET api/v1/health, api/v1/status, api/v1/policy) — api-gateway-regional
    2. regional-api-proxy VPC Lambda: resolves the ALB from SSM //alb-hostname- and signs the HMAC envelope
    3. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    4. manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) — manifest-processor, health-monitor

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
500 Internal error in the proxy Lambda or the backend. —

GET /api/v1/{proxy+}

Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.

Only when: Present only in the commercial aws partition, where Global Accelerator provides the data path to the regional ALBs. Elsewhere the global API keeps only its aggregate routes and workload traffic uses each region's bridge.

  • Operation ID: getApiV1Proxy
  • Tags: control-plane
  • Security: sigv4
  • Integration: aws_proxy, 29 s integration timeout

Backend

  • Lambda: api-gateway-proxy (lambda/api-gateway-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout)
  • Behaviour: Buffered control-plane proxy. Fetches the HMAC signing key from Secrets Manager, rejects region-pinning headers and base64 bodies, signs the method, path, query, body digest, timestamp and nonce into the X-GCO-* envelope and forwards over private-root TLS within 28 s.
  • Then:
    1. AWS Global Accelerator: TCP/443 pass-through to the regional internal ALB
    2. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    3. manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) — manifest-processor, health-monitor

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
403 Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. —
500 Internal error in the proxy Lambda or the backend. —

POST /api/v1/{proxy+}

Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.

Only when: Present only in the commercial aws partition, where Global Accelerator provides the data path to the regional ALBs. Elsewhere the global API keeps only its aggregate routes and workload traffic uses each region's bridge.

  • Operation ID: postApiV1Proxy
  • Tags: control-plane
  • Security: sigv4
  • Integration: aws_proxy, 29 s integration timeout

Backend

  • Lambda: api-gateway-proxy (lambda/api-gateway-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout)
  • Behaviour: Buffered control-plane proxy. Fetches the HMAC signing key from Secrets Manager, rejects region-pinning headers and base64 bodies, signs the method, path, query, body digest, timestamp and nonce into the X-GCO-* envelope and forwards over private-root TLS within 28 s.
  • Then:
    1. AWS Global Accelerator: TCP/443 pass-through to the regional internal ALB
    2. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    3. manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) — manifest-processor, health-monitor

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
403 Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. —
500 Internal error in the proxy Lambda or the backend. —

PUT /api/v1/{proxy+}

Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.

Only when: Present only in the commercial aws partition, where Global Accelerator provides the data path to the regional ALBs. Elsewhere the global API keeps only its aggregate routes and workload traffic uses each region's bridge.

  • Operation ID: putApiV1Proxy
  • Tags: control-plane
  • Security: sigv4
  • Integration: aws_proxy, 29 s integration timeout

Backend

  • Lambda: api-gateway-proxy (lambda/api-gateway-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout)
  • Behaviour: Buffered control-plane proxy. Fetches the HMAC signing key from Secrets Manager, rejects region-pinning headers and base64 bodies, signs the method, path, query, body digest, timestamp and nonce into the X-GCO-* envelope and forwards over private-root TLS within 28 s.
  • Then:
    1. AWS Global Accelerator: TCP/443 pass-through to the regional internal ALB
    2. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    3. manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) — manifest-processor, health-monitor

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
403 Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. —
500 Internal error in the proxy Lambda or the backend. —

PATCH /api/v1/{proxy+}

Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.

Only when: Present only in the commercial aws partition, where Global Accelerator provides the data path to the regional ALBs. Elsewhere the global API keeps only its aggregate routes and workload traffic uses each region's bridge.

  • Operation ID: patchApiV1Proxy
  • Tags: control-plane
  • Security: sigv4
  • Integration: aws_proxy, 29 s integration timeout

Backend

  • Lambda: api-gateway-proxy (lambda/api-gateway-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout)
  • Behaviour: Buffered control-plane proxy. Fetches the HMAC signing key from Secrets Manager, rejects region-pinning headers and base64 bodies, signs the method, path, query, body digest, timestamp and nonce into the X-GCO-* envelope and forwards over private-root TLS within 28 s.
  • Then:
    1. AWS Global Accelerator: TCP/443 pass-through to the regional internal ALB
    2. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    3. manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) — manifest-processor, health-monitor

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
403 Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. —
500 Internal error in the proxy Lambda or the backend. —

DELETE /api/v1/{proxy+}

Catch-all for the control-plane API served by the in-cluster FastAPI services. {proxy+} is the rest of the path — for example jobs, jobs/{namespace}/{name}, manifests, health, status — so the operations available here are the /api/v1/* paths of the manifest-processor and health-monitor documents, minus /api/v1/global/*, which the aggregator serves. The request is forwarded unchanged (the stage prefix is not part of the backend path); the response is returned unchanged.

Only when: Present only in the commercial aws partition, where Global Accelerator provides the data path to the regional ALBs. Elsewhere the global API keeps only its aggregate routes and workload traffic uses each region's bridge.

  • Operation ID: deleteApiV1Proxy
  • Tags: control-plane
  • Security: sigv4
  • Integration: aws_proxy, 29 s integration timeout

Backend

  • Lambda: api-gateway-proxy (lambda/api-gateway-proxy/handler.py; python3.14, handler.lambda_handler, 29 s timeout)
  • Behaviour: Buffered control-plane proxy. Fetches the HMAC signing key from Secrets Manager, rejects region-pinning headers and base64 bodies, signs the method, path, query, body digest, timestamp and nonce into the X-GCO-* envelope and forwards over private-root TLS within 28 s.
  • Then:
    1. AWS Global Accelerator: TCP/443 pass-through to the regional internal ALB
    2. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    3. manifest-processor (catch-all and /api/v1/manifests) or health-monitor (/api/v1/health, /api/v1/metrics, /healthz) — manifest-processor, health-monitor

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
403 Forbidden: the caller's IAM identity or the API's resource policy denied execute-api:Invoke, or the backend rejected the request's HMAC envelope. —
500 Internal error in the proxy Lambda or the backend. —

GET /inference/{proxy+}

Streaming access to deployed inference endpoints. {proxy+} is <endpoint-name> or <endpoint-name>/<serving-path>; the in-cluster inference-proxy resolves the endpoint by name and allows only its serving and health paths (see the inference-proxy document). Bodies are limited to 1 MiB and the connection may stay open for API Gateway's full 15-minute streaming window.

Only when: Present only in the commercial aws partition, where Global Accelerator provides the data path to the regional ALBs. Elsewhere the global API keeps only its aggregate routes and workload traffic uses each region's bridge.

  • Operation ID: getInferenceProxy
  • Tags: inference
  • Security: sigv4
  • Integration: aws_proxy, 900 s integration timeout, response streaming

Backend

  • Lambda: inference-streaming-proxy (lambda/inference-streaming-proxy/index.mjs; nodejs24.x, index.handler, 900 s timeout)
  • Behaviour: Node.js response-streaming proxy (ROUTING_MODE=global). Accepts GET, HEAD and POST under /inference//..., caps the body at 1 MiB, signs the HMAC envelope and streams the model server's response for up to 15 minutes.
  • Then:
    1. AWS Global Accelerator: TCP/443 pass-through to the regional internal ALB
    2. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    3. inference-proxy Service: resolves the endpoint by name, enforces its state and the serving-path allowlist — inference-proxy
    4. The endpoint's model-server Service in gco-inference (, -canary or -proxy), streamed back through every hop

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
404 Not found: the endpoint does not exist in this region or the path is outside the inference serving-path allowlist. —
500 Internal error in the proxy Lambda or the backend. —
502 Bad gateway: the in-cluster proxy or the model server did not answer. —

POST /inference/{proxy+}

Streaming access to deployed inference endpoints. {proxy+} is <endpoint-name> or <endpoint-name>/<serving-path>; the in-cluster inference-proxy resolves the endpoint by name and allows only its serving and health paths (see the inference-proxy document). Bodies are limited to 1 MiB and the connection may stay open for API Gateway's full 15-minute streaming window.

Only when: Present only in the commercial aws partition, where Global Accelerator provides the data path to the regional ALBs. Elsewhere the global API keeps only its aggregate routes and workload traffic uses each region's bridge.

  • Operation ID: postInferenceProxy
  • Tags: inference
  • Security: sigv4
  • Integration: aws_proxy, 900 s integration timeout, response streaming

Backend

  • Lambda: inference-streaming-proxy (lambda/inference-streaming-proxy/index.mjs; nodejs24.x, index.handler, 900 s timeout)
  • Behaviour: Node.js response-streaming proxy (ROUTING_MODE=global). Accepts GET, HEAD and POST under /inference//..., caps the body at 1 MiB, signs the HMAC envelope and streams the model server's response for up to 15 minutes.
  • Then:
    1. AWS Global Accelerator: TCP/443 pass-through to the regional internal ALB
    2. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    3. inference-proxy Service: resolves the endpoint by name, enforces its state and the serving-path allowlist — inference-proxy
    4. The endpoint's model-server Service in gco-inference (, -canary or -proxy), streamed back through every hop

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
404 Not found: the endpoint does not exist in this region or the path is outside the inference serving-path allowlist. —
500 Internal error in the proxy Lambda or the backend. —
502 Bad gateway: the in-cluster proxy or the model server did not answer. —

HEAD /inference/{proxy+}

Streaming access to deployed inference endpoints. {proxy+} is <endpoint-name> or <endpoint-name>/<serving-path>; the in-cluster inference-proxy resolves the endpoint by name and allows only its serving and health paths (see the inference-proxy document). Bodies are limited to 1 MiB and the connection may stay open for API Gateway's full 15-minute streaming window.

Only when: Present only in the commercial aws partition, where Global Accelerator provides the data path to the regional ALBs. Elsewhere the global API keeps only its aggregate routes and workload traffic uses each region's bridge.

  • Operation ID: headInferenceProxy
  • Tags: inference
  • Security: sigv4
  • Integration: aws_proxy, 900 s integration timeout, response streaming

Backend

  • Lambda: inference-streaming-proxy (lambda/inference-streaming-proxy/index.mjs; nodejs24.x, index.handler, 900 s timeout)
  • Behaviour: Node.js response-streaming proxy (ROUTING_MODE=global). Accepts GET, HEAD and POST under /inference//..., caps the body at 1 MiB, signs the HMAC envelope and streams the model server's response for up to 15 minutes.
  • Then:
    1. AWS Global Accelerator: TCP/443 pass-through to the regional internal ALB
    2. Internal ALB — the Kubernetes Gateway gco-system/gco-gateway; HTTPRoute gco-routes picks the Service by longest path prefix — cluster-gateway
    3. inference-proxy Service: resolves the endpoint by name, enforces its state and the serving-path allowlist — inference-proxy
    4. The endpoint's model-server Service in gco-inference (, -canary or -proxy), streamed back through every hop

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
404 Not found: the endpoint does not exist in this region or the path is outside the inference serving-path allowlist. —
500 Internal error in the proxy Lambda or the backend. —
502 Bad gateway: the in-cluster proxy or the model server did not answer. —

GET /studio/callback

Unauthenticated MOCK integration returning an empty 200 body: the Cognito hosted-UI OAuth redirect target. The browser redirect carries the authorization code as a query-string parameter and nothing reads the body.

Only when: Present only when analytics_environment.enabled is true in cdk.json.

  • Operation ID: getStudioCallback
  • Tags: studio
  • Security: none (unauthenticated)
  • Integration: mock

Backend

  • Integration: mock
  • Behaviour: API Gateway answers directly; no Lambda and no backend is involved.

Responses

Status Description Content
200 Success; empty JSON body. —

GET /studio/login

Authenticated with a Cognito ID token from the Studio user pool rather than SigV4. Returns a presigned SageMaker Studio URL for the caller's user profile.

Only when: Present only when analytics_environment.enabled is true in cdk.json.

  • Operation ID: getStudioLogin
  • Tags: studio
  • Security: cognito
  • Integration: aws_proxy, 29 s integration timeout

Backend

  • Lambda: analytics-presigned-url (lambda/analytics-presigned-url/handler.py; python3.13, index.handler)
  • Behaviour: The analytics stack's presigned-URL Lambda (GCOAnalyticsStack): exchanges the Cognito identity for a SageMaker Studio presigned login URL. Synthesized here with an inline stand-in because the analytics stack is not part of this API's template.

Responses

Status Description Content
200 Success; the backend's response is returned unchanged. —
400 Bad request: rejected by the proxy before forwarding (for example an X-GCO-Target-Region header or a base64-encoded body) or by the backend. —
401 Unauthorized: the Cognito ID token is missing, expired or not from the Studio pool. —
404 Not found: the endpoint does not exist in this region or the path is outside the inference serving-path allowlist. —
500 Internal error in the proxy Lambda or the backend. —

Schemas

This api gateway declares no component schemas.