EKS使用预加载机制加速EC2 Nodegroup上大镜像的启动速度

本文介绍EKS中大镜像冷启动问题,通过EventBridge和System Manager实现ECR镜像预加载缓存,将Pod启动时间从110秒优化至2秒。

EKS动手实验合集请参考这里。

EKS 1.36 版本 @2026-09 AWS Global 区域(ap-southeast-1)实测通过,节点操作系统为 Amazon Linux 2023(kubelet v1.36.4-eks-a887778,containerd 2.2.7),测试镜像大小约为 4.1 GiB。

目录

  • 一、背景
  • 二、测试环境准备
  • 三、应用Yaml是否打开缓存开关的对比
  • 四、使用EventBridge构建预加载
  • 五、验证预缓存机制生效
  • 六、参考文档

一、背景

在机器学习等场景下,需要在EKS上运行较大体积的Pod,其Image体积可能达到数个GB乃至数十GB。此时在第一次启动Pod时候,会遇到所谓的冷启动问题,也就是EC2 Nodegroup需要从ECR容器镜像仓库拉取较大尺寸的镜像并解压,然后才能启动Pod。后续启动相同镜像即可利用节点上的缓存,无须重复下载。本文介绍如何通过EventBridge与Systems Manager在镜像发布和节点扩容两个时间点主动预加载镜像,从而消除冷启动时间。

1、优化镜像拉取时间的几种方法

EKS最佳实践文档中提到的几种优化方法如下:

  • 让镜像最小化瘦身;
  • 使用multi-stage builds去除构建阶段才需要的内容;
  • 使用网络带宽与EBS吞吐更高的EC2机型作为Nodegroup;
  • 使用SOCI(Seekable OCI)快照器的并行拉取与解压模式(Parallel Pull mode),通过多连接分段下载和多层并行解压缩短单次拉取时间;
  • 使用Bottlerocket的数据卷或EBS快照预置镜像;
  • 在EC2 Nodegroup上预先缓存镜像,即本文介绍的方法。

以上方法可以组合使用。SOCI并行拉取模式缩短的是每一次拉取的耗时,而本文的预加载方法是把拉取动作提前到Pod调度之前完成,二者解决的是冷启动问题的不同环节。

2、在应用中复用节点上的镜像缓存

在拉起应用的配置文件定义中,增加如下imagePullPolicy标签,即可复用EC2 Nodegroup节点上已经存在的缓存。

    spec:
      containers:
      - name: bigimage1
        image: 133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:1
        imagePullPolicy: IfNotPresent
        ports:
        - containerPort: 80

参数IfNotPresent表示节点上已存在该镜像时直接使用缓存。需要注意,当镜像Tag为latest或者未指定Tag时,Kubernetes默认的拉取策略是Always,每次启动都会访问镜像仓库,因此建议使用明确的版本Tag并显式设置IfNotPresent。对于第一次下载,节点上没有缓存,仍然需要较长时间的冷启动。为了解决冷启动,就需要预加载。

3、通过EventBridge和Systems Manager自动预加载镜像的方案

当新的镜像推送到ECR后,ECR会向EventBridge发送ECR Image Action事件。EventBridge规则匹配该事件后,通过Systems Manager Run Command向集群所有节点下发拉取镜像的命令,实现预加载。对于扩容新增的节点,则通过Systems Manager State Manager的关联(Association)在节点注册到Systems Manager时自动执行相同的拉取命令。整体流程如下:

              推送新版本镜像                        扩容新增节点
                    |                                    |
                    v                                    v
      ECR 发送 ECR Image Action 事件         节点 SSM Agent 注册上线
                    |                                    |
                    v                                    v
      EventBridge 规则匹配 PUSH 事件         State Manager 关联按Tag匹配新节点
                    |                                    |
                    v                                    v
      Run Command 下发到带集群Tag的节点      对新节点执行 AWS-RunShellScript
                    \                                   /
                     v                                 v
         节点执行 ctr -n k8s.io images pull 拉取最新Tag的镜像
                                    |
                                    v
         Pod调度到节点后镜像已存在,跳过拉取直接启动容器

详情参考这篇博客。本文将着重讲述本方案的部署。

4、使用Bottlerocket作为Node底层系统实现快速启动的方案

Bottlerocket是AWS开发的开源的、用于运行容器的操作系统,其体积更小、更安全。EKS 1.36默认使用Amazon Linux 2023作为EC2 Nodegroup的操作系统(Amazon Linux 2的EKS AMI已停止发布),可选使用Bottlerocket来作为操作系统运行容器。二者的区别是,使用Amazon Linux 2023的Nodegroup默认只有一块EBS磁盘,操作系统和容器镜像缓存都在这个磁盘上。而Bottlerocket使用两块EBS磁盘,分别作为OS系统盘和容器镜像的数据盘,可以基于预先拉取好镜像的数据盘EBS快照启动新节点,新节点启动时镜像即已存在。

篇幅所限,本文不会展开描述本方案。详情请参考这篇博客。

5、局限

以上方法均不支持Fargate场景。Fargate的每个Pod运行在独立的计算环境中,没有可复用的节点缓存,Fargate应用拉起无法利用到本文的缓存特性。对于Fargate,可参考文末关于zstd压缩镜像的博客优化拉取速度。

下面开始准备测试环境。

二、测试环境准备

1、构建应用

为了模拟大体积的容器镜像,本文在Amazon Linux 2023的基础镜像上安装Apache,然后在构建阶段通过/dev/urandom生成4个各1 GiB的随机数据文件,每个文件单独形成一个镜像层。随机数据无法被压缩,因此推送到ECR后压缩尺寸依然约为4 GiB,可以真实地反映大镜像的下载耗时。旧版本实验使用的amazonlinux:2基础镜像已于2026年6月结束支持,本文改为amazonlinux:2023。

构建Image的定义如下:

FROM public.ecr.aws/amazonlinux/amazonlinux:2023

# Install apache
RUN dnf install -y httpd \
 && dnf clean all \
 && rm -rf /var/cache/dnf

# Install app
COPY src/run_apache.sh /root/
COPY src/index.html /var/www/html/
RUN chown -R apache:apache /var/www \
 && chmod +x /root/run_apache.sh \
 && echo "ServerName localhost" >> /etc/httpd/conf/httpd.conf

# Simulate a large image: 4 layers x 1 GiB of random data (about 4 GiB in total).
# Random data cannot be compressed, so the compressed size in ECR stays close to 4 GiB.
# Change BUILD_ID on every build to generate brand new layers.
ARG BUILD_ID=1
RUN echo "${BUILD_ID}" > /root/build-id && head -c 1G /dev/urandom > /root/blob-1
RUN echo "${BUILD_ID}" >> /root/build-id && head -c 1G /dev/urandom > /root/blob-2
RUN echo "${BUILD_ID}" >> /root/build-id && head -c 1G /dev/urandom > /root/blob-3
RUN echo "${BUILD_ID}" >> /root/build-id && head -c 1G /dev/urandom > /root/blob-4

EXPOSE 80

# starting script for httpd
CMD ["/bin/bash", "-c", "/root/run_apache.sh"]

其中run_apache.sh的脚本如下:

#!/bin/bash
mkdir -p /run/httpd
/usr/sbin/httpd -D FOREGROUND

src/index.html为任意静态页面即可。以上文件均已放在本仓库的21目录下。构建参数BUILD_ID用于在后续测试中生成全新的镜像层:由于BUILD_ID的值发生变化,其后的RUN指令不会命中构建缓存,会重新生成4个随机数据层,从而模拟发布一个新版本的大镜像。

执行如下命令创建ECR仓库:

aws ecr create-repository --repository-name bigimage --region ap-southeast-1

然后在构建环境上执行如下命令,构建并推送Tag为1的镜像。请替换其中的AWS账户ID:

aws ecr get-login-password --region ap-southeast-1 | docker login --username AWS --password-stdin 133129065110.dkr.ecr.ap-southeast-1.amazonaws.com
docker build --build-arg BUILD_ID=1 -t bigimage:1 .
docker tag bigimage:1 133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:1
docker push 133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:1

本文的构建环境为一台m7i-flex.large的EC2实例(Ubuntu 22.04,Docker 29.8.1),构建耗时166秒,推送耗时35秒。构建完成后查看本地镜像:

IMAGE        ID             DISK USAGE   CONTENT SIZE   EXTRA
bigimage:1   67804948f6fd       8.95GB         4.38GB        

推送完成后,执行如下命令查看ECR上的镜像:

aws ecr describe-images --region ap-southeast-1 --repository-name bigimage \
    --query 'imageDetails[].{tag:imageTags[0],size:imageSizeInBytes,pushed:imagePushedAt,type:imageManifestMediaType}' \
    --output table

返回结果如下:

----------------------------------------------------------------------------------------------------------
|                                             DescribeImages                                             |
+-----------------------------------+-------------+-------+----------------------------------------------+
|              pushed               |    size     |  tag  |                    type                      |
+-----------------------------------+-------------+-------+----------------------------------------------+
|  2026-09-29T20:20:56.595000+08:00 |  1176       |  None |  application/vnd.oci.image.manifest.v1+json  |
|  2026-09-29T20:20:56.579000+08:00 |  4384276568 |  None |  application/vnd.oci.image.manifest.v1+json  |
|  2026-09-29T20:20:56.904000+08:00 |  4384276568 |  1    |  application/vnd.oci.image.index.v1+json     |
+-----------------------------------+-------------+-------+----------------------------------------------+

可以看到,一次推送在ECR上产生了3条记录。这是因为新版本Docker默认使用containerd镜像存储并生成构建证明(Provenance Attestation),推送的是一个OCI镜像索引(Image Index),其中包含一个实际的amd64镜像清单和一个约1 KB的证明清单,只有镜像索引带有Tag。这一点会影响后文查询“最新镜像Tag”的命令写法,后文使用--filter tagStatus=TAGGED排除没有Tag的记录。

注意:如果开发者本机为MacOS,特别是Apple Silicon芯片的Mac,不建议在本机构建推送此类大镜像。一方面ARM架构默认构建出的是arm64镜像,与x86_64节点不匹配;另一方面数GB的上传会占用较长时间。建议使用一台与集群同架构的EC2实例作为构建环境。

2、构建EC2 Nodegroup

如果已经按照实验一创建了名为eksworkshop的集群,则可以跳过本节,直接使用该集群。本文的实测即在实验一的集群上完成,该集群的Nodegroup为3台t3.2xlarge节点(后文输出中的Nodegroup名称为podsubnet-ng,节点的Name标签为eksworkshop-podsubnet-ng-Node)。

如需新建集群,定义如下配置文件,并将其中的三个子网ID替换为实际环境中位于三个可用区的私有子网ID:

apiVersion: eksctl.io/v1alpha5
kind: ClusterConfig

metadata:
  name: eksworkshop
  region: ap-southeast-1
  version: "1.36"

vpc:
  clusterEndpoints:
    publicAccess:  true
    privateAccess: true
  subnets:
    private:
      ap-southeast-1a: { id: subnet-04a7c6e7e1589c953 }
      ap-southeast-1b: { id: subnet-031022a6aab9b9e70 }
      ap-southeast-1c: { id: subnet-0eaf9054aa6daa68e }

kubernetesNetworkConfig:
  serviceIPv4CIDR: 10.50.0.0/24

managedNodeGroups:
  - name: managed-ng
    labels:
      Name: managed-ng
    instanceType: t3.xlarge
    minSize: 2
    desiredCapacity: 2
    maxSize: 4
    privateNetworking: true
    subnets:
      - subnet-04a7c6e7e1589c953
      - subnet-031022a6aab9b9e70
      - subnet-0eaf9054aa6daa68e
    volumeType: gp3
    volumeSize: 100
    volumeIOPS: 3000
    volumeThroughput: 125
    tags:
      nodegroup-name: managed-ng
    iam:
      withAddonPolicies:
        imageBuilder: true
        autoScaler: true
        certManager: true
        efs: true
        ebs: true
        awsLoadBalancerController: true
        xRay: true
        cloudWatch: true

cloudWatch:
  clusterLogging:
    enableTypes: ["api", "audit", "authenticator", "controllerManager", "scheduler"]
    logRetentionInDays: 30

与旧版本配置相比,本配置将version调整为1.36,将已废弃的albIngress替换为awsLoadBalancerController,并将maxSize调整为4,以便后文测试节点扩容。

在以上配置中,使用的是gp3作为EC2 Nodegroup节点组的磁盘,默认为3000 IOPS和125MB/s的吞吐,并且没有额外选配费用,对于一般场景是成本最优选择。对于大镜像,镜像层解压阶段的磁盘写入量约为压缩尺寸的两倍,125MB/s的吞吐往往会成为瓶颈。如果希望加载更快,可酌情上调volumeThroughput与volumeIOPS,但会产生额外的EBS费用。

本方案对节点IAM角色有两项要求,eksctl创建的Nodegroup默认均已满足:

  • AmazonSSMManagedInstanceCore:节点上的SSM Agent注册到Systems Manager所需,eksctl默认会为节点角色附加该策略;
  • ecr:DescribeImages:节点上执行的预加载命令需要查询仓库中最新的镜像Tag。本配置中imageBuilder: true附加的AmazonEC2ContainerRegistryPowerUser包含该权限。如果节点角色只有AmazonEC2ContainerRegistryPullOnly,则需要额外授予ecr:DescribeImages,否则查询Tag的命令会报AccessDenied。

将以上配置文件保存为eks-private-subnet.yaml,然后执行如下命令启动集群。

eksctl create cluster -f eks-private-subnet.yaml

集群启动完成。这个集群将在私有子网启动,包括2个t3.xlarge节点组成NodeGroup。

3、构建测试应用

构建如下应用。为了让3个副本分别调度到3个节点上,配置中增加了topologySpreadConstraints拓扑分布约束:

---
apiVersion: v1
kind: Namespace
metadata: 
  name: bigimage
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: bigimage1
  namespace: bigimage
  labels:
    app: bigimage1
spec:
  replicas: 3
  selector:
    matchLabels:
      app: bigimage1
  template:
    metadata:
      labels:
        app: bigimage1
    spec:
      topologySpreadConstraints:
      - maxSkew: 1
        topologyKey: kubernetes.io/hostname
        whenUnsatisfiable: ScheduleAnyway
        labelSelector:
          matchLabels:
            app: bigimage1
      containers:
      - name: bigimage1
        image: 133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:1
        imagePullPolicy: IfNotPresent
        ports:
        - containerPort: 80
---
apiVersion: v1
kind: Service
metadata:
  name: bigimage1
  namespace: bigimage
  annotations:
    service.beta.kubernetes.io/aws-load-balancer-scheme: internet-facing
    service.beta.kubernetes.io/aws-load-balancer-nlb-target-type: ip
spec:
  loadBalancerClass: service.k8s.aws/nlb
  selector:
    app: bigimage1
  type: LoadBalancer
  ports:
  - protocol: TCP
    port: 80
    targetPort: 80

在以上配置中,imagePullPolicy: IfNotPresent参数表示如果镜像已经存在,则使用缓存;Service的loadBalancerClass: service.k8s.aws/nlb表示由AWS Load Balancer Controller创建NLB,需要集群已按照实验二部署该控制器。将以上文件保存为bigimage.yaml。

下面进行冷启动与缓存启动的对比测试。

三、应用Yaml是否打开缓存开关的对比

1、查询Pod启动时间的脚本

这里参考Github上aws-samples库中的containers-blog-maelstrom/prefetch-data-to-EKSnodes/get-pod-boot-time.sh脚本来统计Pod从被调度(PodScheduled)到就绪(Ready)的时间。原始脚本没有指定Namespace,只能查询default Namespace,因此这里增加了Namespace参数;Pod名称参数改为可选,省略时统计该Namespace下的全部Pod;同时修正了MacOS下按本地时区解析UTC时间的问题。代码如下:

#!/usr/bin/env bash
# Usage: ./get-pod-boot-time.sh <namespace> [pod-name]
# If pod-name is omitted, all pods in the namespace are reported.
namespace="$1"
target_pod="$2"
fail() {
  echo "$@" 1>&2
  exit 1
}
test -n "$namespace" || fail "Usage: $0 <namespace> [pod-name]"
# https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-conditions
get_condition_time() {
  pod="$1"
  condition="$2"
  iso_time=$(kubectl get pod "$pod" --namespace "$namespace" -o json | jq -r ".status.conditions[] | select(.type == \"$condition\" and .status == \"True\") | .lastTransitionTime")
  test -n "$iso_time" || fail "Pod $pod is not in $condition yet"
  if [[ "$(uname)" == "Darwin" ]]; then
      date -j -u -f "%Y-%m-%dT%H:%M:%SZ" "$iso_time" +%s
  else
      date -u -d "$iso_time" +%s
  fi
}
if [[ -n "$target_pod" ]]; then
  pods="$target_pod"
else
  pods=$(kubectl get pods --namespace "$namespace" -o jsonpath='{.items[*].metadata.name}')
fi
for pod in $pods; do
  scheduled_time=$(get_condition_time "$pod" PodScheduled) || exit 1
  ready_time=$(get_condition_time "$pod" Ready) || exit 1
  node=$(kubectl get pod "$pod" --namespace "$namespace" -o jsonpath='{.spec.nodeName}')
  echo "It took approximately $(( ready_time - scheduled_time )) seconds for $pod to boot up on $node"
done

保存到本地,文件名为get-pod-boot-time.sh,并执行chmod +x get-pod-boot-time.sh赋予执行权限。脚本依赖jq,在具有运行kubectl正确权限的环境下,运行命令:

./get-pod-boot-time.sh <namespace> [podname]

即可显示启动时间。

2、首次拉起没有缓存的测试

首先正常启动应用:

kubectl apply -f bigimage.yaml

等待Pod全部变为Running后,执行如下命令查看Pod启动时间:

./get-pod-boot-time.sh bigimage

返回结果如下:

It took approximately 70 seconds for bigimage1-698c8784f4-lxbqt to boot up on ip-192-168-54-190.ap-southeast-1.compute.internal
It took approximately 60 seconds for bigimage1-698c8784f4-v7vpn to boot up on ip-192-168-69-155.ap-southeast-1.compute.internal
It took approximately 64 seconds for bigimage1-698c8784f4-xdvg2 to boot up on ip-192-168-30-108.ap-southeast-1.compute.internal

也可以通过kubelet上报的事件查看拉取镜像的实际耗时:

kubectl get events -n bigimage --field-selector reason=Pulled \
    -o custom-columns=POD:.involvedObject.name,MSG:.message --no-headers

返回结果如下:

bigimage1-698c8784f4-lxbqt   Successfully pulled image "133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:1" in 1m9.12s (1m9.12s including waiting). Image size: 4384279444 bytes.
bigimage1-698c8784f4-v7vpn   Successfully pulled image "133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:1" in 59.146s (59.146s including waiting). Image size: 4384279444 bytes.
bigimage1-698c8784f4-xdvg2   Successfully pulled image "133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:1" in 1m2.913s (1m2.913s including waiting). Image size: 4384279444 bytes.

测试结果为60至70秒,几乎全部时间消耗在镜像拉取上。EC2 Nodegroup节点是t3.2xlarge机型,EBS为100GB的gp3磁盘(3000 IOPS、125MB/s吞吐)。

3、使用缓存的测试

再启动一个新的应用bigimage2,复用相同的镜像。其配置与bigimage1的Deployment部分相同,仅将名称与标签改为bigimage2,完整文件见本仓库的21/bigimage2.yaml。执行如下命令:

kubectl apply -f bigimage2.yaml
./get-pod-boot-time.sh bigimage

返回结果中bigimage2的部分如下:

It took approximately 1 seconds for bigimage2-59c599c547-hvpkf to boot up on ip-192-168-69-155.ap-southeast-1.compute.internal
It took approximately 1 seconds for bigimage2-59c599c547-jg9wr to boot up on ip-192-168-54-190.ap-southeast-1.compute.internal
It took approximately 1 seconds for bigimage2-59c599c547-w4mt6 to boot up on ip-192-168-30-108.ap-southeast-1.compute.internal

对应的事件显示镜像已经存在于节点上,没有发生拉取:

bigimage2-59c599c547-hvpkf   Container image "133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:1" already present on machine and can be accessed by the pod

花费时间约1秒。由此可看出,缓存明显提升了启动速度。

4、测试过程的录屏演示

以下录屏为旧版本实验(EKS 1.28)的录制内容,操作流程与本文一致。

5、测试流程小结

整个流程小结如下:

  • 1、以Amazon Linux 2023容器镜像为基础,安装Apache后体积约为200MB;
  • 2、在构建阶段生成4个各1 GiB的随机数据层,镜像总容量约4.1 GiB;
  • 3、随机数据无法压缩,推送到ECR后压缩尺寸依然为4.1 GiB;
  • 4、在t3.2xlarge机型、100GB gp3磁盘(3000 IOPS、125MB/s吞吐)的节点上,首次拉起应用花费60至70秒;
  • 5、在应用配置文件中指定可利用缓存,使用相同镜像拉起新应用,约1秒启动完毕。

下面配置预加载机制,使首次启动也能命中缓存。

四、使用EventBridge进行预加载

前文介绍过,针对EKS的EC2 Nodegroup,可以使用预加载机制,通过EventBridge触发。在实际使用EKS过程中,有两种场景需要考虑:

  • 在ECR上发布最新版镜像后,现有Nodegroup节点只缓存了旧版本的镜像,需要预加载最新版做缓存。此时方案是 EventBridge + Systems Manager Run Command;
  • EKS的Nodegroup扩容,新增了新的EC2节点,此时新节点上还没有任何一个版本的缓存,需要预加载最新版做缓存。此时方案是 Systems Manager State Manager。

两种场景都通过节点上的eks:cluster-name标签选择目标节点。EKS托管节点组会自动为EC2实例添加eks:cluster-name与eks:nodegroup-name标签,与旧版本实验使用的Name标签相比,无需拼接Nodegroup名称,集群内新增Nodegroup也会自动纳入预加载范围。如果只希望对特定Nodegroup预加载,可将目标改为tag:eks:nodegroup-name。

两种场景的配置过程如下。

1、创建EventBridge要使用的IAM Role

配置IAM Role可以通过AWS控制台,也可以通过AWSCLI,二者效果一样。本文使用CLI创建。

准备IAM Role的信任策略如下:

{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Action": "sts:AssumeRole",
            "Principal": {
                "Service": "events.amazonaws.com"
            },
            "Effect": "Allow",
            "Sid": ""
        }
    ]
}

将其保存为iam-role.json,然后使用AWSCLI如下命令创建这个IAM Role。

aws iam create-role --role-name ecr-push-cache --assume-role-policy-document file://iam-role.json

返回结果如下表示创建成功:

{
    "Role": {
        "Path": "/",
        "RoleName": "ecr-push-cache",
        "RoleId": "AROAR57Y4KKLDNVZCM4TO",
        "Arn": "arn:aws:iam::133129065110:role/ecr-push-cache",
        "CreateDate": "2023-12-26T04:50:16+00:00",
        "AssumeRolePolicyDocument": {
            "Version": "2012-10-17",
            "Statement": [
                {
                    "Action": "sts:AssumeRole",
                    "Principal": {
                        "Service": "events.amazonaws.com"
                    },
                    "Effect": "Allow",
                    "Sid": ""
                }
            ]
        }
    }
}

最终生成的IAM Role名称是ecr-push-cache,后文会使用这个名称。如果之前做过本实验,命令会返回EntityAlreadyExists错误,可以直接复用已有的角色。

2、挂载IAM Policy到上一步创建的Role

编写如下一段IAM Policy,请替换其中的Region代号、AWS账户ID(12位数字),以及将集群名称eksworkshop替换为实际的名称:

{
    "Version": "2012-10-17",
    "Statement": [
        {
            "Sid": "SendCommandToClusterNodes",
            "Action": "ssm:SendCommand",
            "Effect": "Allow",
            "Resource": [
                "arn:aws:ec2:ap-southeast-1:133129065110:instance/*"
            ],
            "Condition": {
                "StringEquals": {
                    "ssm:resourceTag/eks:cluster-name": "eksworkshop"
                }
            }
        },
        {
            "Sid": "UseRunShellScriptDocument",
            "Action": "ssm:SendCommand",
            "Effect": "Allow",
            "Resource": [
                "arn:aws:ssm:ap-southeast-1::document/AWS-RunShellScript"
            ]
        }
    ]
}

以上策略只允许EventBridge对带有eks:cluster-name=eksworkshop标签的实例执行AWS-RunShellScript文档。旧版本实验的策略使用了ec2:ResourceTag/*这种以通配符作为标签键的写法,无法精确限定到某个标签键,本文改为Systems Manager官方文档推荐的ssm:resourceTag/<标签键>条件键。

将其保存为iam-policy.json,然后使用AWSCLI如下命令将这个IAM Policy挂载到上一步创建的IAM Role上。

aws iam put-role-policy --role-name ecr-push-cache --policy-name ecr-push-cache-ssm-policy --policy-document file://iam-policy.json

配置成功则返回命令行,否则会提示错误信息。

3、创建EventBridge规则

配置EventBridge可以通过AWS控制台,也可以通过AWSCLI,二者效果一样。但是由于在AWS控制台上配置EventBridge步骤较多,页面跳转和输入选项复杂,容易遗漏和出错,因此本文通过AWSCLI预先写好的JSON格式的配置文件,快速加载配置。

准备如下一段EventBridge配置规则,替换其中的Name为规则的名称,此处可以自定义。在repository-name位置,输入ECR容器镜像仓库的名称。注意这里不是写完整URI,因此不需要带有容器镜像仓库的网址,也不要带有tag,只是名称。例如本文叫做bigimage。

{
    "Name": "ecr-push-cache",
    "Description": "Rule to trigger SSM Run Command on ECR Image PUSH Action Success",
    "EventPattern": "{\"source\": [\"aws.ecr\"], \"detail-type\": [\"ECR Image Action\"], \"detail\": {\"action-type\": [\"PUSH\"], \"result\": [\"SUCCESS\"], \"repository-name\": [\"bigimage\"]}}"
}

将其保存为event-rule.json,然后使用AWSCLI如下命令将这个规则配置到EventBridge上。请注意文件名和对应region的正确。执行如下命令:

aws events put-rule --cli-input-json file://event-rule.json --region ap-southeast-1

配置成功,返回信息如下:

{
    "RuleArn": "arn:aws:events:ap-southeast-1:133129065110:rule/ecr-push-cache"
}

在配置完成后,通过EventBridge控制台在搜索框内输入名字,即可看到刚才创建的规则。如下截图。

点击规则进入查看详情,可以在第一个标签页Event pattern下看到刚才配置的规则。如下截图。

接下来继续配置规则的Target对象。

4、配置EventBridge规则的Target(用于现有EC2预加载最新版镜像)

准备如下一段Target规则:

{
    "Targets": [{
        "Id": "ecr-push-cache-run-command",
        "Arn": "arn:aws:ssm:ap-southeast-1::document/AWS-RunShellScript",
        "RoleArn": "arn:aws:iam::133129065110:role/ecr-push-cache",
        "Input": "{\"commands\":[\"tag=$(aws ecr describe-images --repository-name bigimage --filter tagStatus=TAGGED --query 'sort_by(imageDetails,& imagePushedAt)[-1].imageTags[0]' --output text)\",\"ctr -n k8s.io images pull --user AWS:$(aws ecr get-login-password) 133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:$tag > /dev/null\",\"ctr -n k8s.io images ls name==133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:$tag\"]}",
        "RunCommandParameters": {
            "RunCommandTargets": [{
                "Key": "tag:eks:cluster-name",
                "Values": ["eksworkshop"]
            }]
        }
    }]
}

Input中的三条命令依次完成如下工作:

  • 查询仓库中最近推送的、带有Tag的镜像的Tag。--filter tagStatus=TAGGED用于排除前文提到的无Tag的镜像清单与证明清单;
  • 使用containerd自带的ctr命令,在Kubernetes使用的k8s.io命名空间中拉取该镜像,只有拉取到该命名空间的镜像才能被kubelet识别为节点缓存。拉取的进度输出较长,这里重定向到/dev/null;
  • 列出拉取完成的镜像,作为命令执行结果便于核查。

节点上的AWS CLI会自动从实例元数据获取区域,因此命令中无需指定--region。Run Command以root身份执行,也无需sudo。

替换的内容如下:

  • Arn部分的Region代号要替换;
  • RoleArn里边的AWS账户ID、IAM Role名称要替换;
  • Input中的ECR仓库名称要替换(多处),AWS账户ID与Region代号要替换(多处);
  • RunCommandTargets中的集群名称要替换(本文是eksworkshop)。

将其保存为rule-target.json,然后使用AWSCLI如下命令将这个规则配置到EventBridge的Rule上。请注意替换上一步使用的Rule规则名称,文件名和对应region的正确。执行如下命令:

aws events put-targets --rule ecr-push-cache --cli-input-json file://rule-target.json --region ap-southeast-1

配置成功,返回如下:

{
    "FailedEntryCount": 0,
    "FailedEntries": []
}

如果之前做过本实验,规则上会残留旧版本Target(Id为Id4000985d-1b4b-4e14-8a45-b04103f9871b,目标为tag:Name)。由于put-targets按Id覆盖,新旧Target会同时存在,应执行如下命令删除旧Target:

aws events list-targets-by-rule --rule ecr-push-cache --region ap-southeast-1
aws events remove-targets --rule ecr-push-cache --ids Id4000985d-1b4b-4e14-8a45-b04103f9871b --region ap-southeast-1

在配置完成后,通过EventBridge控制台,就可以看到规则下能显示出来target了。如下截图。

5、配置Systems Manager的State Manager任务(用于新扩容的EC2预加载最新版镜像)

编辑如下一段配置文件,替换其中Targets下的集群名称,以及commands里边的AWS账户ID、Region代号、ECR仓库名称等。最后一个参数AssociationName是最终在Systems Manager上创建后显示的名称,可自定义输入,后续用于查找和区分任务。

{
  "Name": "AWS-RunShellScript",
  "Targets": [
    {
      "Key": "tag:eks:cluster-name",
      "Values": ["eksworkshop"]
    }
  ],
  "Parameters": {
    "commands": [
      "tag=$(aws ecr describe-images --repository-name bigimage --filter tagStatus=TAGGED --query 'sort_by(imageDetails,& imagePushedAt)[-1].imageTags[0]' --output text)",
      "ctr -n k8s.io images pull --user AWS:$(aws ecr get-login-password) 133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:$tag > /dev/null",
      "ctr -n k8s.io images ls name==133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:$tag"
    ]
  },
  "AssociationName": "ecr-push-image"
}

将其保存为ssm-command.json。如果之前做过本实验,需要先删除以tag:Name为目标的旧关联,否则会并存两个同名关联。执行如下命令查询旧关联的ID并删除:

aws ssm list-associations --region ap-southeast-1 \
    --association-filter-list key=AssociationName,value=ecr-push-image \
    --query 'Associations[].[AssociationId,Targets[0].Key]' --output text
aws ssm delete-association --association-id <旧关联的AssociationId> --region ap-southeast-1

然后执行如下命令创建关联:

aws ssm create-association --cli-input-json file://ssm-command.json --region ap-southeast-1

配置成功则返回:

{
    "AssociationDescription": {
        "Name": "AWS-RunShellScript",
        "AssociationVersion": "1",
        "Date": "2026-09-29T20:27:32.983000+08:00",
        "LastUpdateAssociationDate": "2026-09-29T20:27:32.983000+08:00",
        "Overview": {
            "Status": "Pending",
            "DetailedStatus": "Creating"
        },
        "DocumentVersion": "$DEFAULT",
        "Parameters": {
            "commands": [
                "tag=$(aws ecr describe-images --repository-name bigimage --filter tagStatus=TAGGED --query 'sort_by(imageDetails,& imagePushedAt)[-1].imageTags[0]' --output text)",
                "ctr -n k8s.io images pull --user AWS:$(aws ecr get-login-password) 133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:$tag > /dev/null",
                "ctr -n k8s.io images ls name==133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:$tag"
            ]
        },
        "AssociationId": "9c46f543-b357-4a42-adbf-a1d334f3ea33",
        "Targets": [
            {
                "Key": "tag:eks:cluster-name",
                "Values": [
                    "eksworkshop"
                ]
            }
        ],
        "AssociationName": "ecr-push-image",
        "ApplyOnlyAtCronInterval": false
    }
}

关联创建后会立即在当前所有匹配的节点上执行一次,此后每当有新的匹配节点注册到Systems Manager时再自动执行。执行如下命令查看首次执行结果:

aws ssm describe-association --association-id 9c46f543-b357-4a42-adbf-a1d334f3ea33 \
    --region ap-southeast-1 --query 'AssociationDescription.Overview' --output json

返回结果如下,表示在现有3个节点上执行成功:

{
    "Status": "Success",
    "DetailedStatus": "Success",
    "AssociationStatusAggregatedCount": {
        "Success": 3
    }
}

旧版本实验中,已有缓存的旧节点执行结果会显示为失败。本文使用的containerd 2.x在镜像已经存在时,ctr images pull直接返回成功,因此所有节点均显示为成功。

配置成功后,也可以在控制台上看到这一个任务。从AWS控制台左上角,搜索关键字Systems Manager,点击服务进入。从左侧菜单中,找到Node Tools,点击State Manager,右侧清单中可看到名为ecr-push-image的任务。这个名称是上一步JSON文件中指定的。如下截图。

至此配置完成。现在来验证整个机制工作正常。

五、验证预缓存机制生效

这里验证两个场景:

  • 1、现有EC2 Nodegroup不变,之前启动过旧版本,现在发布新版镜像到ECR,观察所有节点是否会自动获取新镜像作为缓存;
  • 2、集群扩容增加新的EC2节点,观察新节点是否会自动获取最新镜像作为缓存。

1、新发布新版镜像到ECR后查看预加载

在构建环境上执行如下命令,以BUILD_ID=2重新构建镜像,生成4个全新的1 GiB镜像层,并以Tag2推送到ECR:

docker build --build-arg BUILD_ID=2 -t bigimage:2 .
docker tag bigimage:2 133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:2
docker push 133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:2

推送完成后,执行如下命令查看EventBridge触发的Run Command:

aws ssm list-commands --region ap-southeast-1 --max-items 2 \
    --query 'Commands[].[CommandId,RequestedDateTime,Status,TargetCount]' --output text

返回结果如下。镜像索引的推送时间为20:31:10,Run Command在1秒后即被触发,目标节点数为3:

ac8690fc-51db-4366-a23d-e5b0ef1f771c	2026-09-29T20:31:11.327000+08:00	InProgress	3
5039de1e-4929-46a8-acde-350866a221d1	2026-09-29T20:27:33.250000+08:00	Success	3

虽然一次推送在ECR上产生了3条镜像记录,但实测中每次推送只触发了一次Run Command,不会重复拉取。

等待约1分钟后,执行如下命令查看每个节点的执行结果,请替换其中的CommandId:

aws ssm list-command-invocations --region ap-southeast-1 \
    --command-id ac8690fc-51db-4366-a23d-e5b0ef1f771c --details \
    --query 'CommandInvocations[].[InstanceId,Status,CommandPlugins[0].ResponseStartDateTime,CommandPlugins[0].ResponseFinishDateTime]' \
    --output text

返回结果如下,3个节点均在60至72秒内完成了4.1 GiB镜像的预加载:

i-0b02630e9566d7da8	Success	2026-09-29T20:31:11.704000+08:00	2026-09-29T20:32:19.522000+08:00
i-0e7275666a52c20f0	Success	2026-09-29T20:31:11.648000+08:00	2026-09-29T20:32:14.319000+08:00
i-00c1b0a868824b4b2	Success	2026-09-29T20:31:11.689000+08:00	2026-09-29T20:32:23.383000+08:00

如果返回的列表为空,表示规则没有触发或Target配置错误,请检查规则的事件模式、Target的角色权限与节点标签。去掉--query参数可以看到完整的命令输出,其中Output字段为ctr images ls的结果,StandardErrorContent中containerd关于bin_dir配置项即将废弃的告警(DEPRECATION)可以忽略:

REF                                                          TYPE                                    DIGEST                                                                  SIZE    PLATFORMS   LABELS 
133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:2 application/vnd.oci.image.index.v1+json sha256:be0dd7dcde4558de3addaec0e53ebaf8e8285f6999f401aa454b41dc90ebf749 4.1 GiB linux/amd64 -

接下来确认节点上的缓存。旧版本实验通过aws ssm start-session逐台登录节点查询,需要交互式会话。本文改为通过Run Command一次查询全部节点,执行如下命令:

aws ssm send-command --region ap-southeast-1 --document-name AWS-RunShellScript \
    --targets Key=tag:eks:cluster-name,Values=eksworkshop \
    --parameters '{"commands":["ctr -n k8s.io images ls -q 2>/dev/null | grep bigimage"]}' \
    --query Command.CommandId --output text

返回CommandId后,执行如下命令查看结果,请替换其中的CommandId:

aws ssm list-command-invocations --region ap-southeast-1 \
    --command-id 351e06d9-cc65-4128-be7c-4556c8fdfc0b --details \
    --query 'CommandInvocations[].[InstanceId,Status,CommandPlugins[0].Output]' --output text

返回结果如下,每个节点上都同时缓存了Tag1与Tag2两个版本:

i-0b02630e9566d7da8	Success	133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:1
133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:2
133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage@sha256:67804948f6fd5f2b0eb748fd8edf5fca33f7aaeff08bddc0c0d525966fb0506c
i-0e7275666a52c20f0	Success	133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:1
133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:2
133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage@sha256:67804948f6fd5f2b0eb748fd8edf5fca33f7aaeff08bddc0c0d525966fb0506c
i-00c1b0a868824b4b2	Success	133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:1
133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:2
133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage@sha256:67804948f6fd5f2b0eb748fd8edf5fca33f7aaeff08bddc0c0d525966fb0506c

也可以在EC2控制台选中节点,点击Connect按钮,通过Session Manager登录节点后执行sudo ctr -n k8s.io images ls | grep bigimage查询,效果相同。

现在使用Tag2的镜像,在另一个Namespace下启动一个全新的应用。配置文件见本仓库的21/bigimage-new.yaml,其中镜像为bigimage:2、Namespace为bigimage-new。执行如下命令:

kubectl apply -f bigimage-new.yaml
./get-pod-boot-time.sh bigimage-new

返回结果如下:

It took approximately 1 seconds for bigimage-new-946f94d6f-27tcs to boot up on ip-192-168-69-155.ap-southeast-1.compute.internal
It took approximately 1 seconds for bigimage-new-946f94d6f-rbkwv to boot up on ip-192-168-30-108.ap-southeast-1.compute.internal
It took approximately 1 seconds for bigimage-new-946f94d6f-w9bkq to boot up on ip-192-168-54-190.ap-southeast-1.compute.internal

对应事件显示Container image "133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:2" already present on machine。由此可见,新版本镜像从未被任何Pod使用过,但首次启动即命中了预加载的缓存,启动时间由60至70秒缩短到约1秒。

2、集群扩容增加新的EC2节点后查看预加载

在集群扩容场景下,现有应用在现存节点是有缓存的,但Nodegroup中新生成的EC2是没有下载过任何一个镜像的,因此对于这些新节点,所有启动都是冷启动,需要从ECR加载镜像。按照本文配置过Systems Manager State Manager之后,新扩容出来的节点,也会自动缓存最新的镜像。

首先进行节点扩容。请替换Cluster名称、Nodegroup名称和区域代号三个参数,然后执行如下命令。本文环境的Nodegroup名称为podsubnet-ng,由3个节点扩容到4个节点:

eksctl scale nodegroup \
  --cluster eksworkshop \
  --name podsubnet-ng \
  --nodes 4 \
  --nodes-min 3 \
  --nodes-max 6 \
  --region ap-southeast-1

返回结果如下:

2026-09-29 20:33:18 [ℹ]  scaling nodegroup "podsubnet-ng" in cluster eksworkshop
2026-09-29 20:33:21 [ℹ]  initiated scaling of nodegroup
2026-09-29 20:33:21 [ℹ]  to see the status of the scaling run `eksctl get nodegroup --cluster eksworkshop --region ap-southeast-1 --name podsubnet-ng`

新节点在约1分钟后变为Ready。为了确认是否触发了State Manager,执行如下命令查看关联在各节点上的执行情况,请替换其中的AssociationId:

aws ssm describe-association-executions --region ap-southeast-1 \
    --association-id 9c46f543-b357-4a42-adbf-a1d334f3ea33 \
    --query 'AssociationExecutions[0].ExecutionId' --output text
aws ssm describe-association-execution-targets --region ap-southeast-1 \
    --association-id 9c46f543-b357-4a42-adbf-a1d334f3ea33 \
    --execution-id <上一条命令返回的ExecutionId> \
    --query 'AssociationExecutionTargets[].[ResourceId,Status,LastExecutionDate,OutputSource.OutputSourceId]' \
    --output text

返回结果如下。新节点i-05100de21fe76bb2e已被自动纳入关联并执行成功,其余3个节点为创建关联时的执行记录:

i-05100de21fe76bb2e	Success	2026-09-29T20:35:33.549000+08:00	97127484-fc66-4538-93f9-ffb652244fe6
i-00c1b0a868824b4b2	Success	2026-09-29T20:27:36.687000+08:00	5039de1e-4929-46a8-acde-350866a221d1
i-0e7275666a52c20f0	Success	2026-09-29T20:27:35.914000+08:00	5039de1e-4929-46a8-acde-350866a221d1
i-0b02630e9566d7da8	Success	2026-09-29T20:27:35.738000+08:00	5039de1e-4929-46a8-acde-350866a221d1

使用最后一列的CommandId查询新节点的执行详情,可看到预加载从20:34:03开始、20:35:15结束,耗时72秒,拉取的是当前最新的Tag2:

REF                                                          TYPE                                    DIGEST                                                                  SIZE    PLATFORMS   LABELS
133129065110.dkr.ecr.ap-southeast-1.amazonaws.com/bigimage:2 application/vnd.oci.image.index.v1+json sha256:be0dd7dcde4558de3addaec0e53ebaf8e8285f6999f401aa454b41dc90ebf749 4.1 GiB linux/amd64 io.cri-containerd.image=managed

现在将应用扩展到4个副本,使新副本调度到新节点上:

kubectl scale deployment bigimage-new -n bigimage-new --replicas 4
./get-pod-boot-time.sh bigimage-new bigimage-new-946f94d6f-fdzsq

返回结果如下:

It took approximately 1 seconds for bigimage-new-946f94d6f-fdzsq to boot up on ip-192-168-60-172.ap-southeast-1.compute.internal

对应事件同样显示镜像already present on machine。至此验证配置成功。

需要注意,State Manager的执行时机是节点SSM Agent注册上线之后,与kubelet注册到集群、Pod被调度到节点这两个动作是并行的。如果节点刚一Ready就有Pod被调度上来,kubelet会与预加载命令同时拉取同一个镜像,此时Pod的启动时间不会缩短。本方案更适合节点扩容与Pod调度之间存在时间差的场景,例如预先扩容节点以应对计划中的流量高峰。如果需要新节点在Ready时镜像即已存在,可参考第一章提到的Bottlerocket数据卷快照方案。

以下录屏为旧版本实验(EKS 1.28)的录制内容,操作流程与本文一致:

注意:新拉起的节点全新加载一个较大的镜像,耗时主要取决于节点的网络带宽和EBS磁盘吞吐。本文t3.2xlarge配合默认gp3(3000 IOPS、125MB/s吞吐)加载4.1 GiB镜像约需60至72秒。如果EC2 Nodegroup使用网络带宽更高的c/m/r系列机型,并为gp3磁盘配置更高的IOPS与吞吐,或者启用SOCI并行拉取模式,加载时间可进一步缩短。

3、清理环境

实验完成后,执行如下命令删除测试应用与预加载配置。请替换其中的AssociationId与集群参数:

kubectl delete -f bigimage-new.yaml
kubectl delete -f bigimage2.yaml
kubectl delete -f bigimage.yaml
aws ssm delete-association --association-id 9c46f543-b357-4a42-adbf-a1d334f3ea33 --region ap-southeast-1
aws events remove-targets --rule ecr-push-cache --ids ecr-push-cache-run-command --region ap-southeast-1
aws events delete-rule --name ecr-push-cache --region ap-southeast-1
aws iam delete-role-policy --role-name ecr-push-cache --policy-name ecr-push-cache-ssm-policy
aws iam delete-role --role-name ecr-push-cache
eksctl scale nodegroup --cluster eksworkshop --name podsubnet-ng --nodes 3 --nodes-min 3 --nodes-max 6 --region ap-southeast-1

节点上缓存的镜像会占用磁盘空间,kubelet会在磁盘使用率超过镜像回收阈值(默认85%)时自动清理未被使用的镜像。预加载的镜像在被Pod使用之前同样属于“未被使用”的镜像,如果节点磁盘紧张,可能会在Pod调度之前被回收,因此需要为节点预留足够的磁盘空间。ECR仓库中每个4 GiB的镜像版本也会产生持续的存储费用,建议为仓库配置生命周期策略(Lifecycle Policy)只保留最近的若干个版本。

六、参考文档

Start Pods faster by prefetching images

https://aws.amazon.com/blogs/containers/start-pods-faster-by-prefetching-images/

Improve container startup time by caching images

https://github.com/aws-samples/containers-blog-maelstrom/tree/main/prefetch-data-to-EKSnodes

Amazon ECR events and EventBridge

https://docs.aws.amazon.com/AmazonECR/latest/userguide/ecr-eventbridge.html

Restricting Run Command access based on tags

https://docs.aws.amazon.com/systems-manager/latest/userguide/run-command-setting-up.html

Amazon EKS Best Practices - Application Scaling and Performance(AI/ML)

https://docs.aws.amazon.com/eks/latest/best-practices/aiml-performance.html

Introducing Seekable OCI Parallel Pull mode for Amazon EKS

https://aws.amazon.com/blogs/containers/introducing-seekable-oci-parallel-pull-mode-for-amazon-eks/

Reduce container startup time on Amazon EKS with Bottlerocket data volume

https://aws.amazon.com/blogs/containers/reduce-container-startup-time-on-amazon-eks-with-bottlerocket-data-volume/

Reducing AWS Fargate Startup Times with zstd Compressed Container Images

https://aws.amazon.com/blogs/containers/reducing-aws-fargate-startup-times-with-zstd-compressed-container-images/


最后修改于 2026-09-29