跳转到主要内容

1 - 升级 etcd 集群与应用程序

升级 etcd 集群和应用程序的文档列表

本节包含与升级 etcd 集群及应用程序相关的特定文档。

升级策略

升级前请注意,etcd 仅支持以下两种升级场景:

  • 补丁升级:在同一小版本内升级补丁版本(例如 3.7.0 至 3.7.1)。
  • 小版本升级:每次仅升级一个次版本(例如 3.6 至 3.7)。不支持跳过次版本的升级,此类操作很可能失败。请在升级至下一个次版本前,先更新至最新补丁版本。

升级 etcd v3.x 集群

升级至 etcd v2.3

2 - 将 etcd 从 v3.5 升级到 v3.6

升级 etcd 3.5 至 3.6 的流程、检查清单与注意事项

在一般情况下,从 etcd v3.5 升级到 v3.6 可以实现零停机时间的滚动升级:

  • 逐一停止 etcd v3.5 进程,并替换为 etcd v3.6 进程
  • 所有 v3.6 进程运行后,集群即可使用 v3.6 中的新特性

在 开始升级 之前,请通读本指南其余部分以做好准备。

升级检查列表

更新 3.5

重要

升级至 3.6 之前,请确保 所有 3.5 版本的成员均已更新至 3.5.32 或更高版本 。补丁版本 3.5.24 至 3.5.26 修复了多个潜在的升级障碍;3.5.32 增加了 --v2-deprecation=write-only-skip-check 功能,并将 etcdutl check v2store 扩展至可检查 WAL 记录以及 v2 快照。

V2 存储系统

说明

如果未配置 --enable-v2 标志或将其设置为 false,则无需采取进一步操作。

如果配置了 --enable-v2,请运行命令 etcdutl check v2store,以验证 v2store 中是否存在非成员关系(自定义)数据。若不存在自定义数据,可安全移除该标志。否则,请参阅 v2 迁移指南 获取更多详情。

新增标志

+etcd --discovery-token ''
+etcd --discovery-endpoints ''
+etcd --discovery-dial-timeout '2s'
+etcd --discovery-request-timeout '5s'
+etcd --discovery-keepalive-time '2s'
+etcd --discovery-keepalive-timeout '6s'
+etcd --discovery-insecure-transport 'true'
+etcd --discovery-insecure-skip-tls-verify 'false'
+etcd --discovery-cert ''
+etcd --discovery-key ''
+etcd --discovery-cacert ''
+etcd --discovery-user ''
+etcd --discovery-password ''
+etcd --feature-gates
+etcd --log-format

已移除标志

-etcd --enable-v2
-etcd --experimental-enable-v2v3
-etcd --proxy
-etcd --proxy-failure-wait
-etcd --proxy-refresh-interval
-etcd --proxy-dial-timeout
-etcd --proxy-write-timeout
-etcd --proxy-read-timeout

标志已弃用

etcd --experimental-bootstrap-defrag-threshold-megabytes 标志已被弃用.


-etcd --experimental-bootstrap-defrag-threshold-megabytes

+etcd --bootstrap-defrag-threshold-megabytes

etcd --experimental-compaction-batch-limit 标志已被弃用.


-etcd --experimental-compaction-batch-limit

+etcd --compaction-batch-limit

etcd --experimental-compact-hash-check-time 标志已被弃用.


-etcd --experimental-compact-hash-check-time

+etcd --compact-hash-check-time

etcd --experimental-compaction-sleep-interval 标志已被弃用.


-etcd --experimental-compaction-sleep-interval

+etcd --compaction-sleep-interval

etcd --experimental-corrupt-check-time 标志已被弃用.


-etcd --experimental-corrupt-check-time

+etcd --corrupt-check-time

etcd --experimental-enable-distributed-tracing 标志已弃用.


-etcd --experimental-enable-distributed-tracing

+etcd --enable-distributed-tracing

etcd --experimental-distributed-tracing-address 标志已被弃用.


-etcd --experimental-distributed-tracing-address

+etcd --distributed-tracing-address

etcd --experimental-distributed-tracing-instance-id 标志已弃用.


-etcd --experimental-distributed-tracing-instance-id

+etcd --distributed-tracing-instance-id

etcd --experimental-distributed-tracing-sampling-rate 标志已被弃用.


-etcd --experimental-distributed-tracing-sampling-rate

+etcd --distributed-tracing-sampling-rate

etcd --experimental-distributed-tracing-service-name 标志已弃用.


-etcd --experimental-distributed-tracing-service-name

+etcd --distributed-tracing-service-name

etcd --experimental-downgrade-check-time 标志已被弃用.


-etcd --experimental-downgrade-check-time

+etcd --downgrade-check-time

etcd --experimental-max-learners 标志已被弃用.


-etcd --experimental-max-learners

+etcd --max-learners

etcd --experimental-memory-mlock 标志已被弃用.


-etcd --experimental-memory-mlock

+etcd --memory-mlock

etcd --experimental-peer-skip-client-san-verification 标志已被弃用.


-etcd --experimental-peer-skip-client-san-verification

+etcd --peer-skip-client-san-verification

etcd --experimental-snapshot-catchup-entries 标志已被弃用.


-etcd --experimental-snapshot-catchup-entries

+etcd --snapshot-catchup-entries

etcd --experimental-warning-apply-duration 标志已弃用.


-etcd --experimental-warning-apply-duration

+etcd --warning-apply-duration

etcd --experimental-warning-unary-request-duration 标志已被弃用.


-etcd --experimental-warning-unary-request-duration

+etcd --warning-unary-request-duration

etcd --experimental-watch-progress-notify-interval 标志已被弃用.


-etcd --experimental-watch-progress-notify-interval

+etcd --watch-progress-notify-interval

v3.5 功能门控的等效标志

对应的功能门控标志 etcd --experimental-compact-hash-check-enabled=true


-etcd --experimental-compact-hash-check-enabled=true

+etcd --feature-gates=CompactHashCheck=true

对应的功能门控标志 etcd --experimental-initial-corrupt-check=true


-etcd --experimental-initial-corrupt-check=true

+etcd --feature-gates=InitialCorruptCheck=true

对应的功能门控标志 etcd --experimental-enable-lease-checkpoint=true


-etcd --experimental-enable-lease-checkpoint=true

+etcd --feature-gates=LeaseCheckpoint=true

对应的功能门控标志 etcd --experimental-enable-lease-checkpoint-persist=true


-etcd --experimental-enable-lease-checkpoint-persist=true

+etcd --feature-gates=LeaseCheckpointPersist=true

对应的功能门控标志 etcd --experimental-stop-grpc-service-on-defrag=true


-etcd --experimental-stop-grpc-service-on-defrag=true

+etcd --feature-gates=StopGRPCServiceOnDefrag=true

对应的功能门控标志 etcd --experimental-txn-mode-write-with-shared-buffer=false


-etcd --experimental-txn-mode-write-with-shared-buffer=false

+etcd --feature-gates=TxnModeWriteWithSharedBuffer=false

带有新默认值的标志

原始默认标志 etcd --snapshot-count=100000


-etcd --snapshot-count=100000

+etcd --snapshot-count=10000

原始默认标志 etcd --v2-deprecation='not-yet'


-etcd --v2-deprecation='not-yet'

+etcd --v2-deprecation='write-only'

原始默认标志 etcd --discovery-fallback='proxy'


-etcd --discovery-fallback='proxy'

+etcd --discovery-fallback='exit'

Prometheus 指标差异

# metrics added in v3.6
+etcd_network_known_peers
+etcd_server_feature_enabled

服务器升级检查清单

升级要求

要将现有的 etcd 部署升级至 v3.6,运行中的集群版本必须为 v3.5 或更高。若版本低于 v3.5,请先 升级至 v3.5 ,再升级至 v3.6。

此外,为确保滚动升级顺利进行,运行中的集群必须处于健康状态。在继续操作前,请使用 etcdctl endpoint health 命令检查集群健康状况。

准备

在升级 etcd 之前,请务必在预发环境中测试依赖 etcd 的服务,再将升级部署到生产环境。

开始前,请先 下载快照备份 。如果升级出现问题,可以使用此备份 回滚 到现有 etcd 版本。请注意,snapshot 命令只备份 v3 数据。

混合版本

升级期间,etcd 集群支持不同版本的 etcd 成员共存,并以最低公共版本的协议运行。只有当集群中所有成员均升级至 v3.6 版本后,才认为集群已完成升级。内部机制上,etcd 成员之间会相互协商以确定集群的整体版本,该版本控制报告的版本及支持的功能。

回滚

升级 etcd 集群前,请创建并 下载快照备份 。该快照可用于在需要时将集群恢复至升级前的状态。若用户在升级过程中遇到问题,应首先识别并解决根本原因。若集群仍处于混合版本状态——即至少有一个成员仍运行在 v3.5 版本——则可选择将二进制文件或镜像替换为旧版 v3.5,或直接使用快照恢复集群。在此混合状态下,集群仍以 v3.5 集群模式运行,支持回滚而无需执行正式的降级流程。

然而,一旦所有成员均升级至 v3.6 版本,集群即被视为已完全升级,此时使用二进制文件回滚将不再可行。在此情况下,唯一的恢复方式是恢复升级前创建的快照。若用户希望在完成完整升级后返回原始版本,应遵循官方降级指南,以确保一致性并避免数据损坏。

升级流程

本示例演示如何升级在本地计算机上运行的 3 个成员的 v3.5 etcd 集群。

步骤 1: 检查升级要求

集群是否健康且运行 v3.5.x 版本?

etcdctl --endpoints=localhost:2379,localhost:22379,localhost:32379 endpoint health
<<COMMENT
localhost:2379 is healthy: successfully committed proposal: took = 2.555774ms
localhost:32379 is healthy: successfully committed proposal: took = 2.631133ms
localhost:22379 is healthy: successfully committed proposal: took = 3.020958ms
COMMENT

curl http://localhost:2379/version
<<COMMENT
{"etcdserver":"3.5.18","etcdcluster":"3.5.0"}
COMMENT

curl http://localhost:22379/version
<<COMMENT
{"etcdserver":"3.5.18","etcdcluster":"3.5.0"}
COMMENT

curl http://localhost:32379/version
<<COMMENT
{"etcdserver":"3.5.18","etcdcluster":"3.5.0"}
COMMENT

Step 2: 从领导者下载快照备份

下载快照备份 ,以便在出现任何问题时提供回退路径。

etcd 领导者保证拥有最新的应用数据,因此应从领导者获取快照:

curl -sL http://localhost:2379/metrics | grep etcd_server_is_leader
<<COMMENT
# HELP etcd_server_is_leader Whether or not this member is a leader. 1 if is, 0 otherwise.
# TYPE etcd_server_is_leader gauge
etcd_server_is_leader 1
COMMENT

curl -sL http://localhost:22379/metrics | grep etcd_server_is_leader
<<COMMENT
etcd_server_is_leader 0
COMMENT

curl -sL http://localhost:32379/metrics | grep etcd_server_is_leader
<<COMMENT
etcd_server_is_leader 0
COMMENT

etcdctl --endpoints=localhost:2379 snapshot save backup.db
<<COMMENT
{"level":"info","ts":"2025-03-01T04:34:10.336768+0530","caller":"snapshot/v3_snapshot.go:65","msg":"created temporary db file","path":"backup.db.part"}
{"level":"info","ts":"2025-03-01T04:34:10.342373+0530","logger":"client","caller":"v3@v3.5.18/maintenance.go:212","msg":"opened snapshot stream; downloading"}
{"level":"info","ts":"2025-03-01T04:34:10.342433+0530","caller":"snapshot/v3_snapshot.go:73","msg":"fetching snapshot","endpoint":"localhost:2379"}
{"level":"info","ts":"2025-03-01T04:34:10.346482+0530","logger":"client","caller":"v3@v3.5.18/maintenance.go:220","msg":"completed snapshot read; closing"}
{"level":"info","ts":"2025-03-01T04:34:10.348801+0530","caller":"snapshot/v3_snapshot.go:88","msg":"fetched snapshot","endpoint":"localhost:2379","size":"20 kB","took":"now"}
{"level":"info","ts":"2025-03-01T04:34:10.348933+0530","caller":"snapshot/v3_snapshot.go:97","msg":"saved","path":"backup.db"}
Snapshot saved at backup.db
COMMENT

第 3 步:停止一个现有的 etcd 服务器

当每个 etcd 进程停止时,集群中的其他成员会记录预期的错误。这是正常的,因为集群成员之间的连接已(暂时)中断:

{"level":"info","ts":"2025-03-01T04:31:50.654520+0530","caller":"etcdserver/server.go:2676","msg":"cluster version is updated","cluster-version":"3.5"}
{"level":"info","ts":"2025-03-01T04:34:10.345927+0530","caller":"v3rpc/maintenance.go:130","msg":"sending database snapshot to client","total-bytes":20480,"size":"20 kB"}
{"level":"info","ts":"2025-03-01T04:34:10.346094+0530","caller":"v3rpc/maintenance.go:170","msg":"sending database sha256 checksum to client","total-bytes":20480,"checksum-size":32}
{"level":"info","ts":"2025-03-01T04:34:10.346108+0530","caller":"v3rpc/maintenance.go:179","msg":"successfully sent database snapshot to client","total-bytes":20480,"size":"20 kB","took":"now"}
^C
{"level":"info","ts":"2025-03-01T04:35:01.443045+0530","caller":"osutil/interrupt_unix.go:64","msg":"received signal; shutting down","signal":"interrupt"}
{"level":"info","ts":"2025-03-01T04:35:01.443088+0530","caller":"embed/etcd.go:408","msg":"closing etcd server","name":"node1","data-dir":"/tmp/etcd-node1","advertise-peer-urls":["http://127.0.0.1:2380"],"advertise-client-urls":["http://127.0.0.1:2379"]}
{"level":"info","ts":"2025-03-01T04:35:01.443417+0530","caller":"etcdserver/server.go:1503","msg":"leadership transfer starting","local-member-id":"bf9071f4639c75cc","current-leader-member-id":"bf9071f4639c75cc","transferee-member-id":"91bc3c398fb3c146"}
{"level":"info","ts":"2025-03-01T04:35:01.443441+0530","logger":"raft","caller":"etcdserver/zap_raft.go:77","msg":"bf9071f4639c75cc [term 2] starts to transfer leadership to 91bc3c398fb3c146"}
{"level":"info","ts":"2025-03-01T04:35:01.443455+0530","logger":"raft","caller":"etcdserver/zap_raft.go:77","msg":"bf9071f4639c75cc sends MsgTimeoutNow to 91bc3c398fb3c146 immediately as 91bc3c398fb3c146 already has up-to-date log"}
{"level":"warn","ts":"2025-03-01T04:35:01.443517+0530","caller":"embed/serve.go:179","msg":"stopping insecure grpc server due to error","error":"accept tcp 127.0.0.1:2379: use of closed network connection"}
{"level":"warn","ts":"2025-03-01T04:35:01.443548+0530","caller":"embed/serve.go:181","msg":"stopped insecure grpc server due to error","error":"accept tcp 127.0.0.1:2379: use of closed network connection"}
{"level":"info","ts":"2025-03-01T04:35:01.445536+0530","logger":"raft","caller":"etcdserver/zap_raft.go:77","msg":"bf9071f4639c75cc [term: 2] received a MsgVote message with higher term from 91bc3c398fb3c146 [term: 3]"}
{"level":"info","ts":"2025-03-01T04:35:01.445556+0530","logger":"raft","caller":"etcdserver/zap_raft.go:77","msg":"bf9071f4639c75cc became follower at term 3"}
{"level":"info","ts":"2025-03-01T04:35:01.445565+0530","logger":"raft","caller":"etcdserver/zap_raft.go:77","msg":"bf9071f4639c75cc [logterm: 2, index: 12, vote: 0] cast MsgVote for 91bc3c398fb3c146 [logterm: 2, index: 12] at term 3"}
{"level":"info","ts":"2025-03-01T04:35:01.445572+0530","logger":"raft","caller":"etcdserver/zap_raft.go:77","msg":"raft.node: bf9071f4639c75cc lost leader bf9071f4639c75cc at term 3"}
{"level":"info","ts":"2025-03-01T04:35:01.446773+0530","logger":"raft","caller":"etcdserver/zap_raft.go:77","msg":"raft.node: bf9071f4639c75cc elected leader 91bc3c398fb3c146 at term 3"}
{"level":"info","ts":"2025-03-01T04:35:01.544062+0530","caller":"etcdserver/server.go:1520","msg":"leadership transfer finished","local-member-id":"bf9071f4639c75cc","old-leader-member-id":"bf9071f4639c75cc","new-leader-member-id":"91bc3c398fb3c146","took":"100.640374ms"}
{"level":"info","ts":"2025-03-01T04:35:01.544160+0530","caller":"rafthttp/peer.go:330","msg":"stopping remote peer","remote-peer-id":"91bc3c398fb3c146"}
{"level":"warn","ts":"2025-03-01T04:35:01.544956+0530","caller":"rafthttp/stream.go:286","msg":"closed TCP streaming connection with remote peer","stream-writer-type":"stream MsgApp v2","remote-peer-id":"91bc3c398fb3c146"}
{"level":"info","ts":"2025-03-01T04:35:01.544984+0530","caller":"rafthttp/stream.go:294","msg":"stopped TCP streaming connection with remote peer","stream-writer-type":"stream MsgApp v2","remote-peer-id":"91bc3c398fb3c146"}
{"level":"warn","ts":"2025-03-01T04:35:01.545050+0530","caller":"rafthttp/stream.go:286","msg":"closed TCP streaming connection with remote peer","stream-writer-type":"stream Message","remote-peer-id":"91bc3c398fb3c146"}
{"level":"info","ts":"2025-03-01T04:35:01.545065+0530","caller":"rafthttp/stream.go:294","msg":"stopped TCP streaming connection with remote peer","stream-writer-type":"stream Message","remote-peer-id":"91bc3c398fb3c146"}
{"level":"info","ts":"2025-03-01T04:35:01.545099+0530","caller":"rafthttp/pipeline.go:85","msg":"stopped HTTP pipelining with remote peer","local-member-id":"bf9071f4639c75cc","remote-peer-id":"91bc3c398fb3c146"}
{"level":"warn","ts":"2025-03-01T04:35:01.545156+0530","caller":"rafthttp/stream.go:421","msg":"lost TCP streaming connection with remote peer","stream-reader-type":"stream MsgApp v2","local-member-id":"bf9071f4639c75cc","remote-peer-id":"91bc3c398fb3c146","error":"context canceled"}
{"level":"warn","ts":"2025-03-01T04:35:01.545178+0530","caller":"rafthttp/peer_status.go:66","msg":"peer became inactive (message send to peer failed)","peer-id":"91bc3c398fb3c146","error":"failed to read 91bc3c398fb3c146 on stream MsgApp v2 (context canceled)"}
{"level":"info","ts":"2025-03-01T04:35:01.545199+0530","caller":"rafthttp/stream.go:442","msg":"stopped stream reader with remote peer","stream-reader-type":"stream MsgApp v2","local-member-id":"bf9071f4639c75cc","remote-peer-id":"91bc3c398fb3c146"}
{"level":"warn","ts":"2025-03-01T04:35:01.545246+0530","caller":"rafthttp/stream.go:421","msg":"lost TCP streaming connection with remote peer","stream-reader-type":"stream Message","local-member-id":"bf9071f4639c75cc","remote-peer-id":"91bc3c398fb3c146","error":"context canceled"}
{"level":"info","ts":"2025-03-01T04:35:01.545263+0530","caller":"rafthttp/stream.go:442","msg":"stopped stream reader with remote peer","stream-reader-type":"stream Message","local-member-id":"bf9071f4639c75cc","remote-peer-id":"91bc3c398fb3c146"}
{"level":"info","ts":"2025-03-01T04:35:01.545272+0530","caller":"rafthttp/peer.go:335","msg":"stopped remote peer","remote-peer-id":"91bc3c398fb3c146"}
{"level":"info","ts":"2025-03-01T04:35:01.545282+0530","caller":"rafthttp/peer.go:330","msg":"stopping remote peer","remote-peer-id":"fd422379fda50e48"}
{"level":"warn","ts":"2025-03-01T04:35:01.545307+0530","caller":"rafthttp/stream.go:286","msg":"closed TCP streaming connection with remote peer","stream-writer-type":"stream MsgApp v2","remote-peer-id":"fd422379fda50e48"}
{"level":"info","ts":"2025-03-01T04:35:01.545328+0530","caller":"rafthttp/stream.go:294","msg":"stopped TCP streaming connection with remote peer","stream-writer-type":"stream MsgApp v2","remote-peer-id":"fd422379fda50e48"}
{"level":"warn","ts":"2025-03-01T04:35:01.545359+0530","caller":"rafthttp/stream.go:286","msg":"closed TCP streaming connection with remote peer","stream-writer-type":"stream Message","remote-peer-id":"fd422379fda50e48"}
{"level":"info","ts":"2025-03-01T04:35:01.545379+0530","caller":"rafthttp/stream.go:294","msg":"stopped TCP streaming connection with remote peer","stream-writer-type":"stream Message","remote-peer-id":"fd422379fda50e48"}
{"level":"info","ts":"2025-03-01T04:35:01.545410+0530","caller":"rafthttp/pipeline.go:85","msg":"stopped HTTP pipelining with remote peer","local-member-id":"bf9071f4639c75cc","remote-peer-id":"fd422379fda50e48"}
{"level":"warn","ts":"2025-03-01T04:35:01.545467+0530","caller":"rafthttp/stream.go:421","msg":"lost TCP streaming connection with remote peer","stream-reader-type":"stream MsgApp v2","local-member-id":"bf9071f4639c75cc","remote-peer-id":"fd422379fda50e48","error":"context canceled"}
{"level":"warn","ts":"2025-03-01T04:35:01.545485+0530","caller":"rafthttp/peer_status.go:66","msg":"peer became inactive (message send to peer failed)","peer-id":"fd422379fda50e48","error":"failed to read fd422379fda50e48 on stream MsgApp v2 (context canceled)"}
{"level":"info","ts":"2025-03-01T04:35:01.545504+0530","caller":"rafthttp/stream.go:442","msg":"stopped stream reader with remote peer","stream-reader-type":"stream MsgApp v2","local-member-id":"bf9071f4639c75cc","remote-peer-id":"fd422379fda50e48"}
{"level":"warn","ts":"2025-03-01T04:35:01.545560+0530","caller":"rafthttp/stream.go:421","msg":"lost TCP streaming connection with remote peer","stream-reader-type":"stream Message","local-member-id":"bf9071f4639c75cc","remote-peer-id":"fd422379fda50e48","error":"context canceled"}
{"level":"info","ts":"2025-03-01T04:35:01.545577+0530","caller":"rafthttp/stream.go:442","msg":"stopped stream reader with remote peer","stream-reader-type":"stream Message","local-member-id":"bf9071f4639c75cc","remote-peer-id":"fd422379fda50e48"}
{"level":"info","ts":"2025-03-01T04:35:01.545592+0530","caller":"rafthttp/peer.go:335","msg":"stopped remote peer","remote-peer-id":"fd422379fda50e48"}
{"level":"warn","ts":"2025-03-01T04:35:01.545669+0530","caller":"rafthttp/http.go:413","msg":"failed to find remote peer in cluster","local-member-id":"bf9071f4639c75cc","remote-peer-id-stream-handler":"bf9071f4639c75cc","remote-peer-id-from":"91bc3c398fb3c146","cluster-id":"59a05384c9b79ee"}
{"level":"warn","ts":"2025-03-01T04:35:01.545698+0530","caller":"rafthttp/http.go:413","msg":"failed to find remote peer in cluster","local-member-id":"bf9071f4639c75cc","remote-peer-id-stream-handler":"bf9071f4639c75cc","remote-peer-id-from":"fd422379fda50e48","cluster-id":"59a05384c9b79ee"}
{"level":"warn","ts":"2025-03-01T04:35:01.545732+0530","caller":"rafthttp/http.go:413","msg":"failed to find remote peer in cluster","local-member-id":"bf9071f4639c75cc","remote-peer-id-stream-handler":"bf9071f4639c75cc","remote-peer-id-from":"91bc3c398fb3c146","cluster-id":"59a05384c9b79ee"}
{"level":"warn","ts":"2025-03-01T04:35:01.545765+0530","caller":"rafthttp/http.go:413","msg":"failed to find remote peer in cluster","local-member-id":"bf9071f4639c75cc","remote-peer-id-stream-handler":"bf9071f4639c75cc","remote-peer-id-from":"fd422379fda50e48","cluster-id":"59a05384c9b79ee"}
{"level":"info","ts":"2025-03-01T04:35:01.549658+0530","caller":"embed/etcd.go:613","msg":"stopping serving peer traffic","address":"127.0.0.1:2380"}
{"level":"info","ts":"2025-03-01T04:35:02.550532+0530","caller":"embed/etcd.go:618","msg":"stopped serving peer traffic","address":"127.0.0.1:2380"}
{"level":"info","ts":"2025-03-01T04:35:02.550561+0530","caller":"embed/etcd.go:410","msg":"closed etcd server","name":"node1","data-dir":"/tmp/etcd-node1","advertise-peer-urls":["http://127.0.0.1:2380"],"advertise-client-urls":["http://127.0.0.1:2379"]}

第 4 步:使用相同配置重启 etcd 服务器

使用相同配置但采用新 etcd 二进制文件重启 etcd 服务器。

-etcd-old --name ${name} \
+etcd-new --name ${name} \
  --data-dir /path/to/${name}.etcd \
  --listen-client-urls http://localhost:2379 \
  --advertise-client-urls http://localhost:2379 \
  --listen-peer-urls http://localhost:2380 \
  --initial-advertise-peer-urls http://localhost:2380 \
  --initial-cluster s1=http://localhost:2380,s2=http://localhost:22380,s3=http://localhost:32380 \
  --initial-cluster-token tkn \
  --initial-cluster-state new

新的 v3.6 etcd 将向集群发布其信息。此时,集群仍以 v3.5 协议运行,该版本为最低公共版本。

{"level":"info","ts":"2025-03-01T04:40:36.828+0530","caller":"api/capability.go:76","msg":"enabled capabilities for version","cluster-version":"3.5"}

{"level":"info","ts":"2025-03-01T04:40:36.889+0530","caller":"membership/cluster.go:539","msg":"updated cluster version","cluster-id":"59a05384c9b79ee","local-member-id":"bf9071f4639c75cc","from":"3.0","to":"3.5"}

{"level":"info","ts":"2025-03-01T04:40:36.828+0530","caller":"api/capability.go:76","msg":"enabled capabilities for version","cluster-version":"3.5"}

{"level":"info","ts":"2025-03-01T04:40:36.894+0530","caller":"etcdserver/server.go:1686","msg":"published local member to cluster through raft","local-member-id":"bf9071f4639c75cc","local-member-attributes":"{Name:node1 ClientURLs:[http://127.0.0.1:2379]}","cluster-id":"59a05384c9b79ee","publish-timeout":"7s"}

验证每个成员以及整个集群在使用新的 v3.6 etcd 二进制文件后是否恢复正常健康状态:

etcdctl endpoint health --endpoints=localhost:2379,localhost:22379,localhost:32379
<<COMMENT
localhost:2379 is healthy: successfully committed proposal: took = 1.704998ms
localhost:22379 is healthy: successfully committed proposal: took = 2.331754ms
localhost:32379 is healthy: successfully committed proposal: took = 2.490705ms
COMMENT

未升级的成员将持续记录如下警告,直至整个集群完成升级。

这是预期行为,当所有 etcd 集群成员都升级到 v3.6 后,该现象将停止。

{"level":"warn","ts":"2025-03-01T04:40:37.545960+0530","caller":"etcdserver/cluster_util.go:189","msg":"leader found higher-versioned member","local-member-version":"3.5.18","remote-member-id":"bf9071f4639c75cc","remote-member-version":"3.6.0-alpha.0"}

第 5 步:重复第 3 步和第 4 步,对剩余的成员进行操作

所有成员升级完成后,集群将成功报告升级至 v3.6:

成员 1:

{"level":"info","ts":"2025-03-01T04:58:32.375+0530","caller":"etcdserver/server.go:2149","msg":"updating cluster version using v3 API","from":"3.5","to":"3.6"} {"level":"info","ts":"2025-03-01T04:58:32.377+0530","caller":"etcdserver/server.go:2164","msg":"cluster version is updated","cluster-version":"3.6"}

成员 2:

{"level":"info","ts":"2025-03-01T04:58:32.377+0530","caller":"membership/cluster.go:539","msg":"updated cluster version","cluster-id":"59a05384c9b79ee","local-member-id":"91bc3c398fb3c146","from":"3.5","to":"3.6"}

成员 3:

{"level":"info","ts":"2025-03-01T04:58:32.377+0530","caller":"membership/cluster.go:539","msg":"updated cluster version","cluster-id":"59a05384c9b79ee","local-member-id":"fd422379fda50e48","from":"3.5","to":"3.6"}

endpoint health --endpoints=localhost:2379,localhost:22379,localhost:32379
<<COMMENT
localhost:2379 is healthy: successfully committed proposal: took = 492.834µs
localhost:22379 is healthy: successfully committed proposal: took = 1.015025ms
localhost:32379 is healthy: successfully committed proposal: took = 1.853077ms
COMMENT

curl http://localhost:2379/version
<<COMMENT
{"etcdserver":"3.6.0-alpha.0","etcdcluster":"3.6.0"}
COMMENT

curl http://localhost:22379/version
<<COMMENT
{"etcdserver":"3.6.0-alpha.0","etcdcluster":"3.6.0"}
COMMENT

curl http://localhost:32379/version
<<COMMENT
{"etcdserver":"3.6.0-alpha.0","etcdcluster":"3.6.0"}
COMMENT

3 - 将 etcd 从 3.4 升级到 3.5

升级 etcd 3.4 至 3.5 的流程、检查清单与注意事项

在一般情况下,从 etcd 3.4 升级到 3.5 可以实现零停机滚动升级:

  • 逐一停止 etcd v3.4 进程,并替换为 etcd v3.5 进程
  • 在所有 v3.5 进程运行后,集群即可使用 v3.5 的新特性

在 开始升级 之前,请通读本指南其余部分以做好准备。

升级检查列表

警告

从 没有 v3 数据的 v2 迁移 时,如果 etcd 从现有快照恢复,但不存在 v3 ETCD_DATA_DIR/member/snap/db 文件,etcd v3.2+ 服务器会发生崩溃。这种情况出现在服务器由 v2 迁移且此前没有 v3 数据时。此限制也可防止意外丢失 v3 数据(例如 db 文件可能已被移动)。etcd 要求 v3 迁移后的操作必须有 v3 数据。v3.0 服务器包含 v3 数据之前,请勿升级到更新的 v3 版本。

警告

如果集群启用了认证,将无法执行从 3.4 或更早版本的滚动升级,因为 3.5 更改了与认证相关的 WAL 条目格式 。

3.5 版本中的重点变更。

已弃用 etcd_debugging_mvcc_db_total_size_in_bytes Prometheus 指标

v3.5 将 etcd_debugging_mvcc_db_total_size_in_bytes 的 Prometheus 指标提升至 etcd_mvcc_db_total_size_in_bytes,以鼓励对 etcd 存储进行监控。v3.5 完全弃用了 etcd_debugging_mvcc_db_total_size_in_bytes。

-etcd_debugging_mvcc_db_total_size_in_bytes
+etcd_mvcc_db_total_size_in_bytes

请注意,etcd_debugging_* 命名空间指标已被标记为实验性。随着监控指南的完善,我们可能会将更多指标升级为正式支持。

已弃用 etcd_debugging_mvcc_put_total Prometheus 指标

v3.5 将 etcd_debugging_mvcc_put_total 的 Prometheus 指标提升至 etcd_mvcc_put_total,以鼓励对 etcd 存储进行监控。v3.5 完全弃用了 etcd_debugging_mvcc_put_total。

-etcd_debugging_mvcc_put_total
+etcd_mvcc_put_total

请注意,etcd_debugging_* 命名空间指标已被标记为实验性。随着监控指南的完善,我们可能会将更多指标升级为正式支持。

已弃用 etcd_debugging_mvcc_delete_total Prometheus 指标

v3.5 将 etcd_debugging_mvcc_delete_total 的 Prometheus 指标提升至 etcd_mvcc_delete_total,以鼓励对 etcd 存储进行监控。v3.5 完全弃用了 etcd_debugging_mvcc_delete_total。

-etcd_debugging_mvcc_delete_total
+etcd_mvcc_delete_total

请注意,etcd_debugging_* 命名空间指标已被标记为实验性。随着监控指南的完善,我们可能会将更多指标升级为正式支持。

已弃用 etcd_debugging_mvcc_txn_total Prometheus 指标

v3.5 将 etcd_debugging_mvcc_txn_total 的 Prometheus 指标提升至 etcd_mvcc_txn_total,以鼓励对 etcd 存储进行监控。v3.5 完全弃用了 etcd_debugging_mvcc_txn_total。

-etcd_debugging_mvcc_txn_total
+etcd_mvcc_txn_total

请注意,etcd_debugging_* 命名空间指标已被标记为实验性。随着监控指南的完善,我们可能会将更多指标升级为正式支持。

已弃用 etcd_debugging_mvcc_range_total Prometheus 指标

v3.5 将 etcd_debugging_mvcc_range_total Prometheus 指标提升至 etcd_mvcc_range_total 级别,以鼓励对 etcd 存储进行监控。v3.5 完全弃用了 etcd_debugging_mvcc_range_total。

-etcd_debugging_mvcc_range_total
+etcd_mvcc_range_total

请注意,etcd_debugging_* 命名空间指标已被标记为实验性。随着监控指南的完善,我们可能会将更多指标升级为正式支持。

已弃用 etcd --logger capnslog

v3.4 默认使用 --logger=zap,以支持多日志输出和结构化日志。

etcd --logger=capnslog 在 v3.5 中已被弃用,现在 --logger=zap 为默认值。

-etcd --logger=capnslog
+etcd --logger=zap --log-outputs=stderr

+# to write logs to stderr and a.log file at the same time
+etcd --logger=zap --log-outputs=stderr,a.log

v3.4 增加 etcd --logger=zap 对结构化日志和多日志输出的支持。主要动机是推动 etcd 的自动化监控,而非在服务出现异常时回溯服务器日志。未来开发将尽量减少 etcd 的日志输出,并通过指标和告警使 etcd 更易于监控。etcd --logger=capnslog 将在 v3.5 中弃用。

已弃用 etcd --log-output

v3.4 将 etcd --log-output 重命名为 --log-outputs ,以支持多日志输出。

etcd --log-output 已在 v3.5 中弃用.

-etcd --log-output=stderr
+etcd --log-outputs=stderr

已弃用 etcd --debug 标志(现已 --log-level=debug)

etcd --debug 标志已弃用.

-etcd --debug
+etcd --log-level debug

已弃用 etcd --log-package-levels

etcd --log-package-levels 标志在 capnslog 中已弃用.

现在,etcd --logger=zap 为默认值。

-etcd --log-package-levels 'etcdmain=CRITICAL,etcdserver=DEBUG'
+etcd --logger=zap --log-outputs=stderr

已弃用 [CLIENT-URL]/config/local/log

/config/local/log 端点在 v3.5 中正在被弃用,同时 etcd --log-package-levels 标志也将被弃用.

-$ curl http://127.0.0.1:2379/config/local/log -XPUT -d '{"Level":"DEBUG"}'
-# debug logging enabled

变更 gRPC 网关 HTTP 端点(已弃用 /v3beta)

本文未提供内容。

curl -L http://localhost:2379/v3beta/kv/put \
  -X POST -d '{"key": "Zm9v", "value": "YmFy"}'

之后

curl -L http://localhost:2379/v3/kv/put \
  -X POST -d '{"key": "Zm9v", "value": "YmFy"}'

/v3beta 已在 3.5 版本中移除。

服务器升级检查清单

升级要求

要将现有 etcd 部署升级至 3.5 版本,运行中的集群版本必须为 3.4 或更高。若版本低于 3.4,请先 升级至 3.4 ,再升级至 3.5。

此外,为确保滚动升级顺利进行,运行中的集群必须处于健康状态。在继续操作前,请使用 etcdctl endpoint health 命令检查集群健康状况。

准备

在升级 etcd 之前,请务必在预发环境中测试依赖 etcd 的服务,再将升级部署到生产环境。

在开始之前,下载快照备份 。若升级过程中出现异常,可使用此备份将 etcd 版本 回退 至当前版本。请注意,snapshot命令仅备份 v3 数据。如需备份 v2 数据,请参见备份 v2 数据存储 。

混合版本

升级期间,etcd 集群支持不同版本的 etcd 成员共存,并以最低公共版本的协议运行。只有当集群中所有成员均升级至 3.5 版本后,才认为集群已完成升级。内部机制上,etcd 成员之间会相互协商以确定集群的整体版本,该版本控制报告的版本及支持的功能。

限制

请注意:如果集群仅包含 v3 数据且无 v2 数据,则不受此限制影响。

如果集群正在服务的数据集大小超过 50MB,每个新升级的成员可能需要最多 2 分钟才能追上现有集群。请检查最近快照的大小以估算总数据量。换句话说,升级每个成员之间应至少等待 2 分钟。

对于数据总量更大(例如 100MB 或更多)的情况,此一次性操作可能需要更长时间。对于规模达到此类程度的大型 etcd 集群,系统管理员可在升级前自由联系 etcd 团队 ,我们将乐意提供升级流程方面的建议。

降级

如果所有成员均已升级至 v3.5,则集群将升级至 v3.5,从该完成状态回退不可行。然而,若任一成员仍为 v3.4,则集群及其操作仍保持 “v3.4”,在此混合集群状态下,可将所有成员恢复为使用 v3.4 etcd 二进制文件。

请 下载快照备份 ,以便在集群完成升级后仍可执行降级操作。

升级流程

本示例演示如何升级在本地计算机上运行的 3 个成员的 v3.4 etcd 集群。

步骤 1: 检查升级要求

集群是否健康且运行 v3.4.x 版本?

etcdctl --endpoints=localhost:2379,localhost:22379,localhost:32379 endpoint health
<<COMMENT
localhost:2379 is healthy: successfully committed proposal: took = 2.118638ms
localhost:22379 is healthy: successfully committed proposal: took = 3.631388ms
localhost:32379 is healthy: successfully committed proposal: took = 2.157051ms
COMMENT

curl http://localhost:2379/version
<<COMMENT
{"etcdserver":"3.4.0","etcdcluster":"3.4.0"}
COMMENT

curl http://localhost:22379/version
<<COMMENT
{"etcdserver":"3.4.0","etcdcluster":"3.4.0"}
COMMENT

curl http://localhost:32379/version
<<COMMENT
{"etcdserver":"3.4.0","etcdcluster":"3.4.0"}
COMMENT

Step 2: 从领导者下载快照备份

下载快照备份 ,以便在出现任何问题时提供回退路径。

etcd 领导者保证拥有最新的应用数据,因此应从领导者获取快照:

curl -sL http://localhost:2379/metrics | grep etcd_server_is_leader
<<COMMENT
# HELP etcd_server_is_leader Whether or not this member is a leader. 1 if is, 0 otherwise.
# TYPE etcd_server_is_leader gauge
etcd_server_is_leader 1
COMMENT

curl -sL http://localhost:22379/metrics | grep etcd_server_is_leader
<<COMMENT
etcd_server_is_leader 0
COMMENT

curl -sL http://localhost:32379/metrics | grep etcd_server_is_leader
<<COMMENT
etcd_server_is_leader 0
COMMENT

etcdctl --endpoints=localhost:2379 snapshot save backup.db
<<COMMENT
{"level":"info","ts":1526585787.148433,"caller":"snapshot/v3_snapshot.go:109","msg":"created temporary db file","path":"backup.db.part"}
{"level":"info","ts":1526585787.1485257,"caller":"snapshot/v3_snapshot.go:120","msg":"fetching snapshot","endpoint":"localhost:2379"}
{"level":"info","ts":1526585787.1519694,"caller":"snapshot/v3_snapshot.go:133","msg":"fetched snapshot","endpoint":"localhost:2379","took":0.003502721}
{"level":"info","ts":1526585787.1520295,"caller":"snapshot/v3_snapshot.go:142","msg":"saved","path":"backup.db"}
Snapshot saved at backup.db
COMMENT

第 3 步:停止一个现有的 etcd 服务器

当每个 etcd 进程停止时,集群中的其他成员会记录预期的错误。这是正常的,因为集群成员之间的连接已(暂时)中断:

{"level":"info","ts":1526587281.2001143,"caller":"etcdserver/server.go:2249","msg":"updating cluster version","from":"3.0","to":"3.4"}
{"level":"info","ts":1526587281.2010646,"caller":"membership/cluster.go:473","msg":"updated cluster version","cluster-id":"7dee9ba76d59ed53","local-member-id":"7339c4e5e833c029","from":"3.0","from":"3.4"}
{"level":"info","ts":1526587281.2012327,"caller":"api/capability.go:76","msg":"enabled capabilities for version","cluster-version":"3.4"}
{"level":"info","ts":1526587281.2013083,"caller":"etcdserver/server.go:2272","msg":"cluster version is updated","cluster-version":"3.4"}



^C{"level":"info","ts":1526587299.0717514,"caller":"osutil/interrupt_unix.go:63","msg":"received signal; shutting down","signal":"interrupt"}
{"level":"info","ts":1526587299.0718873,"caller":"embed/etcd.go:285","msg":"closing etcd server","name":"s1","data-dir":"/tmp/etcd/s1","advertise-peer-urls":["http://localhost:2380"],"advertise-client-urls":["http://localhost:2379"]}
{"level":"info","ts":1526587299.0722554,"caller":"etcdserver/server.go:1341","msg":"leadership transfer starting","local-member-id":"7339c4e5e833c029","current-leader-member-id":"7339c4e5e833c029","transferee-member-id":"729934363faa4a24"}
{"level":"info","ts":1526587299.0723994,"caller":"raft/raft.go:1107","msg":"7339c4e5e833c029 [term 3] starts to transfer leadership to 729934363faa4a24"}
{"level":"info","ts":1526587299.0724802,"caller":"raft/raft.go:1113","msg":"7339c4e5e833c029 sends MsgTimeoutNow to 729934363faa4a24 immediately as 729934363faa4a24 already has up-to-date log"}
{"level":"info","ts":1526587299.0737045,"caller":"raft/raft.go:797","msg":"7339c4e5e833c029 [term: 3] received a MsgVote message with higher term from 729934363faa4a24 [term: 4]"}
{"level":"info","ts":1526587299.0737681,"caller":"raft/raft.go:656","msg":"7339c4e5e833c029 became follower at term 4"}
{"level":"info","ts":1526587299.073831,"caller":"raft/raft.go:882","msg":"7339c4e5e833c029 [logterm: 3, index: 9, vote: 0] cast MsgVote for 729934363faa4a24 [logterm: 3, index: 9] at term 4"}
{"level":"info","ts":1526587299.0738947,"caller":"raft/node.go:312","msg":"raft.node: 7339c4e5e833c029 lost leader 7339c4e5e833c029 at term 4"}
{"level":"info","ts":1526587299.0748374,"caller":"raft/node.go:306","msg":"raft.node: 7339c4e5e833c029 elected leader 729934363faa4a24 at term 4"}
{"level":"info","ts":1526587299.1726425,"caller":"etcdserver/server.go:1362","msg":"leadership transfer finished","local-member-id":"7339c4e5e833c029","old-leader-member-id":"7339c4e5e833c029","new-leader-member-id":"729934363faa4a24","took":0.100389359}
{"level":"info","ts":1526587299.1728148,"caller":"rafthttp/peer.go:333","msg":"stopping remote peer","remote-peer-id":"b548c2511513015"}
{"level":"warn","ts":1526587299.1751974,"caller":"rafthttp/stream.go:291","msg":"closed TCP streaming connection with remote peer","stream-writer-type":"stream MsgApp v2","remote-peer-id":"b548c2511513015"}
{"level":"warn","ts":1526587299.1752589,"caller":"rafthttp/stream.go:301","msg":"stopped TCP streaming connection with remote peer","stream-writer-type":"stream MsgApp v2","remote-peer-id":"b548c2511513015"}
{"level":"warn","ts":1526587299.177348,"caller":"rafthttp/stream.go:291","msg":"closed TCP streaming connection with remote peer","stream-writer-type":"stream Message","remote-peer-id":"b548c2511513015"}
{"level":"warn","ts":1526587299.1774004,"caller":"rafthttp/stream.go:301","msg":"stopped TCP streaming connection with remote peer","stream-writer-type":"stream Message","remote-peer-id":"b548c2511513015"}
{"level":"info","ts":1526587299.177515,"caller":"rafthttp/pipeline.go:86","msg":"stopped HTTP pipelining with remote peer","local-member-id":"7339c4e5e833c029","remote-peer-id":"b548c2511513015"}
{"level":"warn","ts":1526587299.1777067,"caller":"rafthttp/stream.go:436","msg":"lost TCP streaming connection with remote peer","stream-reader-type":"stream MsgApp v2","local-member-id":"7339c4e5e833c029","remote-peer-id":"b548c2511513015","error":"read tcp 127.0.0.1:34636->127.0.0.1:32380: use of closed network connection"}
{"level":"info","ts":1526587299.1778402,"caller":"rafthttp/stream.go:459","msg":"stopped stream reader with remote peer","stream-reader-type":"stream MsgApp v2","local-member-id":"7339c4e5e833c029","remote-peer-id":"b548c2511513015"}
{"level":"warn","ts":1526587299.1780295,"caller":"rafthttp/stream.go:436","msg":"lost TCP streaming connection with remote peer","stream-reader-type":"stream Message","local-member-id":"7339c4e5e833c029","remote-peer-id":"b548c2511513015","error":"read tcp 127.0.0.1:34634->127.0.0.1:32380: use of closed network connection"}
{"level":"info","ts":1526587299.1780987,"caller":"rafthttp/stream.go:459","msg":"stopped stream reader with remote peer","stream-reader-type":"stream Message","local-member-id":"7339c4e5e833c029","remote-peer-id":"b548c2511513015"}
{"level":"info","ts":1526587299.1781602,"caller":"rafthttp/peer.go:340","msg":"stopped remote peer","remote-peer-id":"b548c2511513015"}
{"level":"info","ts":1526587299.1781986,"caller":"rafthttp/peer.go:333","msg":"stopping remote peer","remote-peer-id":"729934363faa4a24"}
{"level":"warn","ts":1526587299.1802843,"caller":"rafthttp/stream.go:291","msg":"closed TCP streaming connection with remote peer","stream-writer-type":"stream MsgApp v2","remote-peer-id":"729934363faa4a24"}
{"level":"warn","ts":1526587299.1803446,"caller":"rafthttp/stream.go:301","msg":"stopped TCP streaming connection with remote peer","stream-writer-type":"stream MsgApp v2","remote-peer-id":"729934363faa4a24"}
{"level":"warn","ts":1526587299.1824749,"caller":"rafthttp/stream.go:291","msg":"closed TCP streaming connection with remote peer","stream-writer-type":"stream Message","remote-peer-id":"729934363faa4a24"}
{"level":"warn","ts":1526587299.18255,"caller":"rafthttp/stream.go:301","msg":"stopped TCP streaming connection with remote peer","stream-writer-type":"stream Message","remote-peer-id":"729934363faa4a24"}
{"level":"info","ts":1526587299.18261,"caller":"rafthttp/pipeline.go:86","msg":"stopped HTTP pipelining with remote peer","local-member-id":"7339c4e5e833c029","remote-peer-id":"729934363faa4a24"}
{"level":"warn","ts":1526587299.1827736,"caller":"rafthttp/stream.go:436","msg":"lost TCP streaming connection with remote peer","stream-reader-type":"stream MsgApp v2","local-member-id":"7339c4e5e833c029","remote-peer-id":"729934363faa4a24","error":"read tcp 127.0.0.1:51482->127.0.0.1:22380: use of closed network connection"}
{"level":"info","ts":1526587299.182845,"caller":"rafthttp/stream.go:459","msg":"stopped stream reader with remote peer","stream-reader-type":"stream MsgApp v2","local-member-id":"7339c4e5e833c029","remote-peer-id":"729934363faa4a24"}
{"level":"warn","ts":1526587299.1830168,"caller":"rafthttp/stream.go:436","msg":"lost TCP streaming connection with remote peer","stream-reader-type":"stream Message","local-member-id":"7339c4e5e833c029","remote-peer-id":"729934363faa4a24","error":"context canceled"}
{"level":"warn","ts":1526587299.1831107,"caller":"rafthttp/peer_status.go:65","msg":"peer became inactive","peer-id":"729934363faa4a24","error":"failed to read 729934363faa4a24 on stream Message (context canceled)"}
{"level":"info","ts":1526587299.1831737,"caller":"rafthttp/stream.go:459","msg":"stopped stream reader with remote peer","stream-reader-type":"stream Message","local-member-id":"7339c4e5e833c029","remote-peer-id":"729934363faa4a24"}
{"level":"info","ts":1526587299.1832306,"caller":"rafthttp/peer.go:340","msg":"stopped remote peer","remote-peer-id":"729934363faa4a24"}
{"level":"warn","ts":1526587299.1837125,"caller":"rafthttp/http.go:424","msg":"failed to find remote peer in cluster","local-member-id":"7339c4e5e833c029","remote-peer-id-stream-handler":"7339c4e5e833c029","remote-peer-id-from":"b548c2511513015","cluster-id":"7dee9ba76d59ed53"}
{"level":"warn","ts":1526587299.1840093,"caller":"rafthttp/http.go:424","msg":"failed to find remote peer in cluster","local-member-id":"7339c4e5e833c029","remote-peer-id-stream-handler":"7339c4e5e833c029","remote-peer-id-from":"b548c2511513015","cluster-id":"7dee9ba76d59ed53"}
{"level":"warn","ts":1526587299.1842315,"caller":"rafthttp/http.go:424","msg":"failed to find remote peer in cluster","local-member-id":"7339c4e5e833c029","remote-peer-id-stream-handler":"7339c4e5e833c029","remote-peer-id-from":"729934363faa4a24","cluster-id":"7dee9ba76d59ed53"}
{"level":"warn","ts":1526587299.1844475,"caller":"rafthttp/http.go:424","msg":"failed to find remote peer in cluster","local-member-id":"7339c4e5e833c029","remote-peer-id-stream-handler":"7339c4e5e833c029","remote-peer-id-from":"729934363faa4a24","cluster-id":"7dee9ba76d59ed53"}
{"level":"info","ts":1526587299.2056687,"caller":"embed/etcd.go:473","msg":"stopping serving peer traffic","address":"127.0.0.1:2380"}
{"level":"info","ts":1526587299.205819,"caller":"embed/etcd.go:480","msg":"stopped serving peer traffic","address":"127.0.0.1:2380"}
{"level":"info","ts":1526587299.2058413,"caller":"embed/etcd.go:289","msg":"closed etcd server","name":"s1","data-dir":"/tmp/etcd/s1","advertise-peer-urls":["http://localhost:2380"],"advertise-client-urls":["http://localhost:2379"]}

第 4 步:使用相同配置重启 etcd 服务器

使用相同配置但采用新 etcd 二进制文件重启 etcd 服务器。

-etcd-old --name s1 \
+etcd-new --name s1 \
  --data-dir /tmp/etcd/s1 \
  --listen-client-urls http://localhost:2379 \
  --advertise-client-urls http://localhost:2379 \
  --listen-peer-urls http://localhost:2380 \
  --initial-advertise-peer-urls http://localhost:2380 \
  --initial-cluster s1=http://localhost:2380,s2=http://localhost:22380,s3=http://localhost:32380 \
  --initial-cluster-token tkn \
  --initial-cluster-state new

新的 v3.5 etcd 将向集群发布其信息。此时,集群仍以 v3.4 协议运行,该版本为最低公共版本。

{"level":"info","ts":1526586617.1647713,"caller":"membership/cluster.go:485","msg":"set initial cluster version","cluster-id":"7dee9ba76d59ed53","local-member-id":"7339c4e5e833c029","cluster-version":"3.0"}

{"level":"info","ts":1526586617.1648536,"caller":"api/capability.go:76","msg":"enabled capabilities for version","cluster-version":"3.0"}

{"level":"info","ts":1526586617.1649303,"caller":"membership/cluster.go:473","msg":"updated cluster version","cluster-id":"7dee9ba76d59ed53","local-member-id":"7339c4e5e833c029","from":"3.0","from":"3.4"}

{"level":"info","ts":1526586617.1649797,"caller":"api/capability.go:76","msg":"enabled capabilities for version","cluster-version":"3.4"}

{"level":"info","ts":1526586617.2107732,"caller":"etcdserver/server.go:1770","msg":"published local member to cluster through raft","local-member-id":"7339c4e5e833c029","local-member-attributes":"{Name:s1 ClientURLs:[http://localhost:2379]}","request-path":"/0/members/7339c4e5e833c029/attributes","cluster-id":"7dee9ba76d59ed53","publish-timeout":7}

验证每个成员以及整个集群在使用新的 v3.5 etcd 二进制文件后是否恢复正常健康状态:

etcdctl endpoint health --endpoints=localhost:2379,localhost:22379,localhost:32379
<<COMMENT
localhost:32379 is healthy: successfully committed proposal: took = 2.337471ms
localhost:22379 is healthy: successfully committed proposal: took = 1.130717ms
localhost:2379 is healthy: successfully committed proposal: took = 2.124843ms
COMMENT

未升级的成员将持续记录如下警告,直至整个集群完成升级。

这是预期行为,当所有 etcd 集群成员都升级到 v3.5 后,该现象将停止。

:41.942121 W | etcdserver: member 7339c4e5e833c029 has a higher version 3.5.0
:45.945154 W | etcdserver: the local etcd version 3.4.0 is not up-to-date

第 5 步:重复第 3 步和第 4 步,对剩余的成员进行操作

所有成员升级完成后,集群将成功报告升级至 3.5:

成员 1:

{"level":"info","ts":1526586949.0920913,"caller":"api/capability.go:76","msg":"enabled capabilities for version","cluster-version":"3.5"} {"level":"info","ts":1526586949.0921566,"caller":"etcdserver/server.go:2272","msg":"cluster version is updated","cluster-version":"3.5"}

成员 2:

{"level":"info","ts":1526586949.092117,"caller":"membership/cluster.go:473","msg":"updated cluster version","cluster-id":"7dee9ba76d59ed53","local-member-id":"729934363faa4a24","from":"3.4","from":"3.5"} {"level":"info","ts":1526586949.0923078,"caller":"api/capability.go:76","msg":"enabled capabilities for version","cluster-version":"3.5"}

成员 3:

{"level":"info","ts":1526586949.0921423,"caller":"membership/cluster.go:473","msg":"updated cluster version","cluster-id":"7dee9ba76d59ed53","local-member-id":"b548c2511513015","from":"3.4","from":"3.5"} {"level":"info","ts":1526586949.0922918,"caller":"api/capability.go:76","msg":"enabled capabilities for version","cluster-version":"3.5"}

endpoint health --endpoints=localhost:2379,localhost:22379,localhost:32379
<<COMMENT
localhost:2379 is healthy: successfully committed proposal: took = 492.834µs
localhost:22379 is healthy: successfully committed proposal: took = 1.015025ms
localhost:32379 is healthy: successfully committed proposal: took = 1.853077ms
COMMENT

curl http://localhost:2379/version
<<COMMENT
{"etcdserver":"3.5.0","etcdcluster":"3.5.0"}
COMMENT

curl http://localhost:22379/version
<<COMMENT
{"etcdserver":"3.5.0","etcdcluster":"3.5.0"}
COMMENT

curl http://localhost:32379/version
<<COMMENT
{"etcdserver":"3.5.0","etcdcluster":"3.5.0"}
COMMENT

4 - 将 etcd 从 3.3 升级到 3.4

升级 etcd 3.3 至 3.4 的流程、检查清单与注意事项

在一般情况下,从 etcd 3.3 升级到 3.4 可以实现零停机滚动升级:

  • 逐一停止 etcd v3.3 进程,并替换为 etcd v3.4 进程
  • 在所有 v3.4 进程运行后,集群即可使用 v3.4 的新特性

在 开始升级 之前,请通读本指南其余部分以做好准备。

升级检查列表

警告

从 没有 v3 数据的 v2 迁移 时,如果 etcd 从现有快照恢复,但不存在 v3 ETCD_DATA_DIR/member/snap/db 文件,etcd v3.2+ 服务器会发生崩溃。这种情况出现在服务器由 v2 迁移且此前没有 v3 数据时。此限制也可防止意外丢失 v3 数据(例如 db 文件可能已被移动)。etcd 要求 v3 迁移后的操作必须有 v3 数据。v3.0 服务器包含 v3 数据之前,请勿升级到更新的 v3 版本。

3.4 版本中的重点变更。

设置 ETCDCTL_API=3 etcdctl 为默认值

ETCDCTL_API=3 现为默认值。

etcdctl set foo bar
Error: unknown command "set" for "etcdctl"

-etcdctl set foo bar
+ETCDCTL_API=2 etcdctl set foo bar
bar

ETCDCTL_API=3 etcdctl put foo bar
OK

-ETCDCTL_API=3 etcdctl put foo bar
+etcdctl put foo bar

设置 etcd --enable-v2=false 为默认值

etcd --enable-v2=false 现为默认值。

这意味着,除非指定了 etcd --enable-v2=true,否则 etcd v3.4 服务器将不会提供 v2 API 请求服务。

如果使用了 v2 API,请确保在 v3.4 版本中启用了 v2 API:

-etcd
+etcd --enable-v2=true

其他 HTTP API 仍可正常工作(例如 [CLIENT-URL]/metrics、[CLIENT-URL]/health、v3 gRPC 网关)。

已弃用 etcd --ca-file 和 etcd --peer-ca-file 标志

--ca-file 和 --peer-ca-file 标志已弃用;自 v2.1 版本起已弃用。

请注意,设置此参数将自动启用客户端证书身份认证,无论 --client-cert-auth 设置为何值。

-etcd --ca-file ca-client.crt
+etcd --trusted-ca-file ca-client.crt
-etcd --peer-ca-file ca-peer.crt
+etcd --peer-trusted-ca-file ca-peer.crt

废弃的grpc.ErrClientConnClosing错误

grpc.ErrClientConnClosing 在 gRPC ≥ 1.10 中已被 弃用 。

import (
+	"go.etcd.io/etcd/clientv3"

	"google.golang.org/grpc"
+	"google.golang.org/grpc/codes"
+	"google.golang.org/grpc/status"
)

_, err := kvc.Get(ctx, "a")
-if err == grpc.ErrClientConnClosing {
+if clientv3.IsConnCanceled(err) {

// or
+s, ok := status.FromError(err)
+if ok {
+  if s.Code() == codes.Canceled

要求 grpc.WithBlock 进行客户端连接

新的客户端负载均衡器 使用异步解析器,将端点传递给 gRPC 连接函数。因此,v3.4 客户端必须使用 grpc.WithBlock 连接选项,以等待底层连接建立完成。

import (
	"time"
	"go.etcd.io/etcd/clientv3"
+	"google.golang.org/grpc"
)

+// "grpc.WithBlock()" to block until the underlying connection is up
ccfg := clientv3.Config{
  Endpoints:            []string{"localhost:2379"},
  DialTimeout:          time.Second,
+ DialOptions:          []grpc.DialOption{grpc.WithBlock()},
  DialKeepAliveTime:    time.Second,
  DialKeepAliveTimeout: 500 * time.Millisecond,
}

废弃 etcd_debugging_mvcc_db_total_size_in_bytes Prometheus 指标

v3.4 将 etcd_debugging_mvcc_db_total_size_in_bytes Prometheus 指标提升至 etcd_mvcc_db_total_size_in_bytes,以鼓励对 etcd 存储进行监控。

etcd_debugging_mvcc_db_total_size_in_bytes 在 v3.4 版本中仍为向后兼容而提供,将在 v3.5 版本中完全弃用。

-etcd_debugging_mvcc_db_total_size_in_bytes
+etcd_mvcc_db_total_size_in_bytes

请注意,etcd_debugging_* 命名空间指标已被标记为实验性。随着监控指南的完善,我们可能会将更多指标升级为正式支持。

废弃 etcd_debugging_mvcc_put_total Prometheus 指标

v3.4 将 etcd_debugging_mvcc_put_total Prometheus 指标提升至 etcd_mvcc_put_total,以鼓励对 etcd 存储进行监控。

etcd_debugging_mvcc_put_total 在 v3.4 版本中仍为向后兼容而提供,将在 v3.5 版本中完全弃用。

-etcd_debugging_mvcc_put_total
+etcd_mvcc_put_total

请注意,etcd_debugging_* 命名空间指标已被标记为实验性。随着监控指南的完善,我们可能会将更多指标升级为正式支持。

废弃 etcd_debugging_mvcc_delete_total Prometheus 指标

v3.4 将 etcd_debugging_mvcc_delete_total Prometheus 指标提升至 etcd_mvcc_delete_total,以鼓励对 etcd 存储进行监控。

etcd_debugging_mvcc_delete_total 在 v3.4 版本中仍为向后兼容而提供,将在 v3.5 版本中完全弃用。

-etcd_debugging_mvcc_delete_total
+etcd_mvcc_delete_total

请注意,etcd_debugging_* 命名空间指标已被标记为实验性。随着监控指南的完善,我们可能会将更多指标升级为正式支持。

废弃 etcd_debugging_mvcc_txn_total Prometheus 指标

v3.4 将 etcd_debugging_mvcc_txn_total Prometheus 指标提升至 etcd_mvcc_txn_total,以鼓励对 etcd 存储进行监控。

etcd_debugging_mvcc_txn_total 在 v3.4 版本中仍为向后兼容而提供,将在 v3.5 版本中完全弃用。

-etcd_debugging_mvcc_txn_total
+etcd_mvcc_txn_total

请注意,etcd_debugging_* 命名空间指标已被标记为实验性。随着监控指南的完善,我们可能会将更多指标升级为正式支持。

废弃 Prometheus 元度指标etcd_debugging_mvcc_range_total

v3.4 将 etcd_debugging_mvcc_range_total 的 Prometheus 指标提升至 etcd_mvcc_range_total,以鼓励对 etcd 存储进行监控。

etcd_debugging_mvcc_range_total 在 v3.4 版本中仍为向后兼容而提供,将在 v3.5 版本中完全弃用。

-etcd_debugging_mvcc_range_total
+etcd_mvcc_range_total

请注意,etcd_debugging_* 命名空间指标已被标记为实验性。随着监控指南的完善,我们可能会将更多指标升级为正式支持。

弃用 etcd --log-output 标志(现已 --log-outputs)

将 etcd --log-output 重命名为 --log-outputs ,以支持多日志输出。etcd --logger=capnslog 不支持多日志输出。

etcd --log-output 将在 v3.5 版本中弃用。etcd --logger=capnslog 将在 v3.5 版本中弃用。

-etcd --log-output=stderr
+etcd --log-outputs=stderr

+# to write logs to stderr and a.log file at the same time
+# only "--logger=zap" supports multiple writers
+etcd --logger=zap --log-outputs=stderr,a.log

v3.4 增加 etcd --logger=zap --log-outputs=stderr 对结构化日志和多日志输出的支持。主要动机是推动 etcd 的自动化监控,而非在服务出现异常时回溯服务器日志。未来开发将尽量减少 etcd 的日志输出,并通过指标和告警使 etcd 更易于监控。etcd --logger=capnslog 将在 v3.5 中弃用。

将 log-outputs 字段类型在 etcd --config-file 中更改为 []string

现在 log-outputs(旧字段名 log-output)支持多个写入者,因此 etcd 配置 YAML 文件 log-outputs 字段必须更改为如下所示的 []string 类型:

 # Specify 'stdout' or 'stderr' to skip journald logging even when running under systemd.
-log-output: default
+log-outputs: [default]

将embed.Config.LogOutput重命名为embed.Config.LogOutputs

将 embed.Config.LogOutput 重命名为 embed.Config.LogOutputs ,以支持多日志输出。并将 embed.Config.LogOutput 类型从 string 改为 []string ,以支持多日志输出。

import "github.com/coreos/etcd/embed"

cfg := &embed.Config{Debug: false}
-cfg.LogOutput = "stderr"
+cfg.LogOutputs = []string{"stderr"}

v3.5 弃用capnslog

v3.5 将弃用 etcd --log-package-levels 标志的 capnslog 功能;etcd --logger=zap --log-outputs=stderr 将成为默认值。v3.5 将弃用 [CLIENT-URL]/config/local/log 端点。

-etcd
+etcd --logger zap

弃用 etcd --debug 标志(现已 --log-level=debug)

v3.4 已弃用 etcd --debug 标志。应改用 etcd --log-level=debug 标志。

-etcd --debug
+etcd --logger zap --log-level debug

弃用的 pkg/transport.TLSInfo.CAFile 字段

已弃用 pkg/transport.TLSInfo.CAFile 字段。

import "github.com/coreos/etcd/pkg/transport"

tlsInfo := transport.TLSInfo{
    CertFile: "/tmp/test-certs/test.pem",
    KeyFile: "/tmp/test-certs/test-key.pem",
-   CAFile: "/tmp/test-certs/trusted-ca.pem",
+   TrustedCAFile: "/tmp/test-certs/trusted-ca.pem",
}
tlsConfig, err := tlsInfo.ClientConfig()
if err != nil {
    panic(err)
}

将 embed.Config.SnapCount 更改为 embed.Config.SnapshotCount

为与标志名称 etcd --snapshot-count 保持一致,embed.Config.SnapCount 字段已重命名为 embed.Config.SnapshotCount:

import "github.com/coreos/etcd/embed"

cfg := embed.NewConfig()
-cfg.SnapCount = 100000
+cfg.SnapshotCount = 100000

将 etcdserver.ServerConfig.SnapCount 更改为 etcdserver.ServerConfig.SnapshotCount

为与标志名称 etcd --snapshot-count 保持一致,etcdserver.ServerConfig.SnapCount 字段已重命名为 etcdserver.ServerConfig.SnapshotCount:

import "github.com/coreos/etcd/etcdserver"

srvcfg := etcdserver.ServerConfig{
-  SnapCount: 100000,
+  SnapshotCount: 100000,

修改了包 wal 的函数签名

修改 wal 函数签名以支持结构化日志记录。

import "github.com/coreos/etcd/wal"
+import "go.uber.org/zap"

+lg, _ = zap.NewProduction()

-wal.Open(dirpath, snap)
+wal.Open(lg, dirpath, snap)

-wal.OpenForRead(dirpath, snap)
+wal.OpenForRead(lg, dirpath, snap)

-wal.Repair(dirpath)
+wal.Repair(lg, dirpath)

-wal.Create(dirpath, metadata)
+wal.Create(lg, dirpath, metadata)

更改了 IntervalTree 类型 在 pkg/adt 包中

pkg/adt.IntervalTree 现已定义为 interface。

import (
    "fmt"

    "go.etcd.io/etcd/pkg/adt"
)

func main() {
-    ivt := &adt.IntervalTree{}
+    ivt := adt.NewIntervalTree()

已弃用 embed.Config.SetupLogging

embed.Config.SetupLogging 已被移除,以防止错误的日志配置,现在将自动设置。

import "github.com/coreos/etcd/embed"

cfg := &embed.Config{Debug: false}
-cfg.SetupLogging()

Changed gRPC 网关 HTTP 端点(替换 /v3beta 为 /v3)

本文未提供内容。

curl -L http://localhost:2379/v3beta/kv/put \
  -X POST -d '{"key": "Zm9v", "value": "YmFy"}'

之后

curl -L http://localhost:2379/v3/kv/put \
  -X POST -d '{"key": "Zm9v", "value": "YmFy"}'

对 /v3beta 端点的请求将重定向至 /v3,/v3beta 将在 3.5 版本中移除。

已弃用的容器镜像标签

latest 及其小版本镜像标签已弃用:

-docker pull gcr.io/etcd-development/etcd:latest
+docker pull gcr.io/etcd-development/etcd:v3.4.0

-docker pull gcr.io/etcd-development/etcd:v3.4
+docker pull gcr.io/etcd-development/etcd:v3.4.0

-docker pull gcr.io/etcd-development/etcd:v3.4
+docker pull gcr.io/etcd-development/etcd:v3.4.1

-docker pull gcr.io/etcd-development/etcd:v3.4
+docker pull gcr.io/etcd-development/etcd:v3.4.2

服务器升级检查清单

升级要求

要将现有 etcd 部署升级至 3.4 版本,运行中的集群版本必须为 3.3 或更高。若版本低于 3.3,请先 升级至 3.3 ,再升级至 3.4。

此外,为确保滚动升级顺利进行,运行中的集群必须处于健康状态。在继续操作前,请使用 etcdctl endpoint health 命令检查集群健康状况。

准备

在升级 etcd 之前,请务必在预发环境中测试依赖 etcd 的服务,再将升级部署到生产环境。

在开始之前,下载快照备份 。若升级过程中出现异常,可使用此备份将 etcd 版本 回退 至当前版本。请注意,snapshot命令仅备份 v3 数据。如需备份 v2 数据,请参见备份 v2 数据存储 。

混合版本

升级期间,etcd 集群支持不同版本的 etcd 成员共存,并以最低公共版本的协议运行。只有当集群中所有成员均升级至 3.4 版本后,该集群才被视为已完成升级。内部机制上,etcd 成员之间会相互协商以确定集群的整体版本,该版本控制报告的版本及支持的功能。

限制

请注意:如果集群仅包含 v3 数据且无 v2 数据,则不受此限制影响。

如果集群正在服务的数据集大小超过 50MB,每个新升级的成员可能需要最多 2 分钟才能追上现有集群。请检查最近快照的大小以估算总数据量。换句话说,升级每个成员之间应至少等待 2 分钟。

对于数据总量更大(例如 100MB 或更多)的情况,此一次性操作可能需要更长时间。对于规模达到此类程度的大型 etcd 集群,系统管理员可在升级前自由联系 etcd 团队 ,我们将乐意提供升级流程方面的建议。

降级

如果所有成员均已升级至 v3.4 版本,集群将升级至 v3.4 版本,从该完成状态回退不可行。然而,若任一成员仍为 v3.3 版本,则集群及其操作仍保持 “v3.3” 状态,此时可从该混合集群状态恢复至所有成员均使用 v3.3 etcd 二进制文件。

请 下载快照备份 ,以便在集群完成升级后仍可执行降级操作。

升级流程

本示例演示如何升级在本地计算机上运行的 3 个成员的 v3.3 etcd 集群。

步骤 1: 检查升级要求

集群是否健康且运行 v3.3.x 版本?

etcdctl --endpoints=localhost:2379,localhost:22379,localhost:32379 endpoint health
<<COMMENT
localhost:2379 is healthy: successfully committed proposal: took = 2.118638ms
localhost:22379 is healthy: successfully committed proposal: took = 3.631388ms
localhost:32379 is healthy: successfully committed proposal: took = 2.157051ms
COMMENT

curl http://localhost:2379/version
<<COMMENT
{"etcdserver":"3.3.5","etcdcluster":"3.3.0"}
COMMENT

curl http://localhost:22379/version
<<COMMENT
{"etcdserver":"3.3.5","etcdcluster":"3.3.0"}
COMMENT

curl http://localhost:32379/version
<<COMMENT
{"etcdserver":"3.3.5","etcdcluster":"3.3.0"}
COMMENT

Step 2: 从领导者下载快照备份

下载快照备份 ,以便在出现任何问题时提供回退路径。

etcd 领导者保证拥有最新的应用数据,因此应从领导者获取快照:

curl -sL http://localhost:2379/metrics | grep etcd_server_is_leader
<<COMMENT
# HELP etcd_server_is_leader Whether or not this member is a leader. 1 if is, 0 otherwise.
# TYPE etcd_server_is_leader gauge
etcd_server_is_leader 1
COMMENT

curl -sL http://localhost:22379/metrics | grep etcd_server_is_leader
<<COMMENT
etcd_server_is_leader 0
COMMENT

curl -sL http://localhost:32379/metrics | grep etcd_server_is_leader
<<COMMENT
etcd_server_is_leader 0
COMMENT

etcdctl --endpoints=localhost:2379 snapshot save backup.db
<<COMMENT
{"level":"info","ts":1526585787.148433,"caller":"snapshot/v3_snapshot.go:109","msg":"created temporary db file","path":"backup.db.part"}
{"level":"info","ts":1526585787.1485257,"caller":"snapshot/v3_snapshot.go:120","msg":"fetching snapshot","endpoint":"localhost:2379"}
{"level":"info","ts":1526585787.1519694,"caller":"snapshot/v3_snapshot.go:133","msg":"fetched snapshot","endpoint":"localhost:2379","took":0.003502721}
{"level":"info","ts":1526585787.1520295,"caller":"snapshot/v3_snapshot.go:142","msg":"saved","path":"backup.db"}
Snapshot saved at backup.db
COMMENT

第 3 步:停止一个现有的 etcd 服务器

当每个 etcd 进程停止时,集群中的其他成员会记录预期的错误。这是正常的,因为集群成员之间的连接已(暂时)中断:

10.237579 I | etcdserver: updating the cluster version from 3.0 to 3.3
10.238315 N | etcdserver/membership: updated the cluster version from 3.0 to 3.3
10.238451 I | etcdserver/api: enabled capabilities for version 3.3


^C21.192174 N | pkg/osutil: received interrupt signal, shutting down...
21.192459 I | etcdserver: 7339c4e5e833c029 starts leadership transfer from 7339c4e5e833c029 to 729934363faa4a24
21.192569 I | raft: 7339c4e5e833c029 [term 8] starts to transfer leadership to 729934363faa4a24
21.192619 I | raft: 7339c4e5e833c029 sends MsgTimeoutNow to 729934363faa4a24 immediately as 729934363faa4a24 already has up-to-date log
WARNING: 2018/05/17 12:45:21 grpc: addrConn.resetTransport failed to create client transport: connection error: desc = "transport: Error while dialing dial tcp: operation was canceled"; Reconnecting to {localhost:2379 0  <nil>}
WARNING: 2018/05/17 12:45:21 grpc: addrConn.transportMonitor exits due to: grpc: the connection is closing
21.193589 I | raft: 7339c4e5e833c029 [term: 8] received a MsgVote message with higher term from 729934363faa4a24 [term: 9]
21.193626 I | raft: 7339c4e5e833c029 became follower at term 9
21.193651 I | raft: 7339c4e5e833c029 [logterm: 8, index: 9, vote: 0] cast MsgVote for 729934363faa4a24 [logterm: 8, index: 9] at term 9
21.193675 I | raft: raft.node: 7339c4e5e833c029 lost leader 7339c4e5e833c029 at term 9
21.194424 I | raft: raft.node: 7339c4e5e833c029 elected leader 729934363faa4a24 at term 9
21.292898 I | etcdserver: 7339c4e5e833c029 finished leadership transfer from 7339c4e5e833c029 to 729934363faa4a24 (took 100.436391ms)
21.292975 I | rafthttp: stopping peer 729934363faa4a24...
21.293206 I | rafthttp: closed the TCP streaming connection with peer 729934363faa4a24 (stream MsgApp v2 writer)
21.293225 I | rafthttp: stopped streaming with peer 729934363faa4a24 (writer)
21.293437 I | rafthttp: closed the TCP streaming connection with peer 729934363faa4a24 (stream Message writer)
21.293459 I | rafthttp: stopped streaming with peer 729934363faa4a24 (writer)
21.293514 I | rafthttp: stopped HTTP pipelining with peer 729934363faa4a24
21.293590 W | rafthttp: lost the TCP streaming connection with peer 729934363faa4a24 (stream MsgApp v2 reader)
21.293610 I | rafthttp: stopped streaming with peer 729934363faa4a24 (stream MsgApp v2 reader)
21.293680 W | rafthttp: lost the TCP streaming connection with peer 729934363faa4a24 (stream Message reader)
21.293700 I | rafthttp: stopped streaming with peer 729934363faa4a24 (stream Message reader)
21.293711 I | rafthttp: stopped peer 729934363faa4a24
21.293720 I | rafthttp: stopping peer b548c2511513015...
21.293987 I | rafthttp: closed the TCP streaming connection with peer b548c2511513015 (stream MsgApp v2 writer)
21.294063 I | rafthttp: stopped streaming with peer b548c2511513015 (writer)
21.294467 I | rafthttp: closed the TCP streaming connection with peer b548c2511513015 (stream Message writer)
21.294561 I | rafthttp: stopped streaming with peer b548c2511513015 (writer)
21.294742 I | rafthttp: stopped HTTP pipelining with peer b548c2511513015
21.294867 W | rafthttp: lost the TCP streaming connection with peer b548c2511513015 (stream MsgApp v2 reader)
21.294892 I | rafthttp: stopped streaming with peer b548c2511513015 (stream MsgApp v2 reader)
21.294990 W | rafthttp: lost the TCP streaming connection with peer b548c2511513015 (stream Message reader)
21.295004 E | rafthttp: failed to read b548c2511513015 on stream Message (context canceled)
21.295013 I | rafthttp: peer b548c2511513015 became inactive
21.295024 I | rafthttp: stopped streaming with peer b548c2511513015 (stream Message reader)
21.295035 I | rafthttp: stopped peer b548c2511513015

第 4 步:使用相同配置重启 etcd 服务器

使用相同配置但采用新 etcd 二进制文件重启 etcd 服务器。

-etcd-old --name s1 \
+etcd-new --name s1 \
  --data-dir /tmp/etcd/s1 \
  --listen-client-urls http://localhost:2379 \
  --advertise-client-urls http://localhost:2379 \
  --listen-peer-urls http://localhost:2380 \
  --initial-advertise-peer-urls http://localhost:2380 \
  --initial-cluster s1=http://localhost:2380,s2=http://localhost:22380,s3=http://localhost:32380 \
  --initial-cluster-token tkn \
+ --initial-cluster-state new \
+ --logger zap \
+ --log-outputs stderr

新的 v3.4 etcd 将向集群发布其信息。此时,集群仍以 v3.3 协议运行,该版本为最低公共版本。

{"level":"info","ts":1526586617.1647713,"caller":"membership/cluster.go:485","msg":"set initial cluster version","cluster-id":"7dee9ba76d59ed53","local-member-id":"7339c4e5e833c029","cluster-version":"3.0"}

{"level":"info","ts":1526586617.1648536,"caller":"api/capability.go:76","msg":"enabled capabilities for version","cluster-version":"3.0"}

{"level":"info","ts":1526586617.1649303,"caller":"membership/cluster.go:473","msg":"updated cluster version","cluster-id":"7dee9ba76d59ed53","local-member-id":"7339c4e5e833c029","from":"3.0","from":"3.3"}

{"level":"info","ts":1526586617.1649797,"caller":"api/capability.go:76","msg":"enabled capabilities for version","cluster-version":"3.3"}

{"level":"info","ts":1526586617.2107732,"caller":"etcdserver/server.go:1770","msg":"published local member to cluster through raft","local-member-id":"7339c4e5e833c029","local-member-attributes":"{Name:s1 ClientURLs:[http://localhost:2379]}","request-path":"/0/members/7339c4e5e833c029/attributes","cluster-id":"7dee9ba76d59ed53","publish-timeout":7}

验证每个成员以及整个集群在使用新的 v3.4 etcd 二进制文件后是否恢复正常健康状态:

etcdctl endpoint health --endpoints=localhost:2379,localhost:22379,localhost:32379
<<COMMENT
localhost:32379 is healthy: successfully committed proposal: took = 2.337471ms
localhost:22379 is healthy: successfully committed proposal: took = 1.130717ms
localhost:2379 is healthy: successfully committed proposal: took = 2.124843ms
COMMENT

未升级的成员将持续记录如下警告,直至整个集群完成升级。

这是预期行为,当所有 etcd 集群成员都升级到 v3.4 后,该现象将停止。

:41.942121 W | etcdserver: member 7339c4e5e833c029 has a higher version 3.4.0
:45.945154 W | etcdserver: the local etcd version 3.3.5 is not up-to-date

第 5 步:重复第 3 步和第 4 步,对剩余的成员进行操作

所有成员升级完成后,集群将成功报告升级至 3.4:

成员 1:

{"level":"info","ts":1526586949.0920913,"caller":"api/capability.go:76","msg":"enabled capabilities for version","cluster-version":"3.4"} {"level":"info","ts":1526586949.0921566,"caller":"etcdserver/server.go:2272","msg":"cluster version is updated","cluster-version":"3.4"}

成员 2:

{"level":"info","ts":1526586949.092117,"caller":"membership/cluster.go:473","msg":"updated cluster version","cluster-id":"7dee9ba76d59ed53","local-member-id":"729934363faa4a24","from":"3.3","from":"3.4"} {"level":"info","ts":1526586949.0923078,"caller":"api/capability.go:76","msg":"enabled capabilities for version","cluster-version":"3.4"}

成员 3:

{"level":"info","ts":1526586949.0921423,"caller":"membership/cluster.go:473","msg":"updated cluster version","cluster-id":"7dee9ba76d59ed53","local-member-id":"b548c2511513015","from":"3.3","from":"3.4"} {"level":"info","ts":1526586949.0922918,"caller":"api/capability.go:76","msg":"enabled capabilities for version","cluster-version":"3.4"}

endpoint health --endpoints=localhost:2379,localhost:22379,localhost:32379
<<COMMENT
localhost:2379 is healthy: successfully committed proposal: took = 492.834µs
localhost:22379 is healthy: successfully committed proposal: took = 1.015025ms
localhost:32379 is healthy: successfully committed proposal: took = 1.853077ms
COMMENT

curl http://localhost:2379/version
<<COMMENT
{"etcdserver":"3.4.0","etcdcluster":"3.4.0"}
COMMENT

curl http://localhost:22379/version
<<COMMENT
{"etcdserver":"3.4.0","etcdcluster":"3.4.0"}
COMMENT

curl http://localhost:32379/version
<<COMMENT
{"etcdserver":"3.4.0","etcdcluster":"3.4.0"}
COMMENT

5 - 将 etcd 从 v3.6 升级到 v3.7

升级 etcd 3.6 至 3.7 的流程、检查清单与注意事项

在一般情况下,从 etcd v3.6 升级到 v3.7 可以实现零停机滚动升级:

  • 逐一停止 etcd v3.6 进程,并替换为 etcd v3.7 进程
  • 在所有 v3.7 进程运行后,集群即可使用 v3.7 中的新特性

在 开始升级 之前,请通读本指南其余部分以做好准备。

升级检查列表

更新 3.6

重要

在升级到 3.7 之前,请确保所有 3.6 成员均已更新至 3.6.11 或更高版本。较早的 3.6 补丁版本可能与 3.7 的滚动升级不兼容。

V2 存储系统

v3.7 版本中已完全移除 v2 存储。v2 HTTP API(--enable-v2)、v2-on-v3 模拟层(--experimental-enable-v2v3)、v2 发现服务、client/v2 包,以及 v2 快照文件的加载功能均已不可用。请参阅 CHANGELOG-3.7 中的破坏性变更说明。

如果从 3.6 版本集群升级,这些标志已不存在,无需采取任何操作。如果从包含自定义 v2 数据的旧版本升级,请在升级前遵循 v2 迁移指南 。

Go 重构

v3.7 包含重大的内部重构,对正常升级流程无影响,但在升级自定义集成时值得留意:

  • 从 gogo/protobuf 迁移到标准 google.golang.org/protobuf(跟踪于 #14533 )。
  • 已将已弃用的 go-grpc-middleware v1 日志和标签库迁移至 v2 拦截器(#20420 )。
  • OpenTelemetry gRPC 拦截器已更新至 otelgrpc v0.61.0,用 NewServerHandler 替代已弃用的 UnaryServerInterceptor 和 StreamServerInterceptor(#20017 )。

如果将 etcd 作为库嵌入,或针对 clientv3 API 进行构建,或依赖内部包,请在升级前查阅 CHANGELOG 。

已移除标志

v3.7 版本已移除所有已弃用的 --experimental-* 标志(#19959 )。在 v3.6 版本中,这些标志均已被同名的非实验性标志或 --feature-gates 条目替代。如果仍存在这些标志的设置,请务必在升级至 v3.7 之前,将其替换为 v3.6 对应的等效设置,否则 v3.7 进程将无法启动。

-etcd --experimental-bootstrap-defrag-threshold-megabytes
-etcd --experimental-compact-hash-check-enabled
-etcd --experimental-compact-hash-check-time
-etcd --experimental-compaction-batch-limit
-etcd --experimental-compaction-sleep-interval
-etcd --experimental-corrupt-check-time
-etcd --experimental-distributed-tracing-address
-etcd --experimental-distributed-tracing-instance-id
-etcd --experimental-distributed-tracing-sampling-rate
-etcd --experimental-distributed-tracing-service-name
-etcd --experimental-downgrade-check-time
-etcd --experimental-enable-distributed-tracing
-etcd --experimental-enable-lease-checkpoint
-etcd --experimental-enable-lease-checkpoint-persist
-etcd --experimental-initial-corrupt-check
-etcd --experimental-memory-mlock
-etcd --experimental-peer-skip-client-san-verification
-etcd --experimental-snapshot-catchup-entries
-etcd --experimental-stop-grpc-service-on-defrag
-etcd --experimental-txn-mode-write-with-shared-buffer
-etcd --experimental-warning-apply-duration
-etcd --experimental-warning-unary-request-duration
-etcd --experimental-watch-progress-notify-interval

请参阅 v3.5 到 v3.6 升级指南 ,以获取每个已移除标志与其非实验性等效标志的映射关系,或查阅 --feature-gates 条目。

新增标志

None.

带有新默认值的标志

None.

服务器升级检查清单

升级要求

要将现有 etcd 部署升级至 v3.7,运行中的集群必须为 v3.6.11 或更高版本。若当前版本为较旧的小版本,请先 升级至 v3.6 ;etcd 仅支持一次升级一个次要版本。

此外,为确保滚动升级顺利进行,运行中的集群必须处于健康状态。在继续操作前,请使用 etcdctl endpoint health 命令检查集群健康状况。

准备

在升级 etcd 之前,请务必在预发环境中测试依赖 etcd 的服务,再将升级部署到生产环境。

开始之前,下载快照备份 。若升级过程中出现异常,可使用此备份 回滚 至现有 etcd 版本。

混合版本

升级期间,etcd 集群支持不同版本的 etcd 成员共存,并以最低共同版本的协议运行。只有当集群中所有成员均升级至 v3.7 版本后,该集群才被视为已完成升级。内部机制上,etcd 成员之间会相互协商以确定集群的整体版本,该版本控制报告的版本及所支持的功能。

回滚

升级 etcd 集群前,请创建并 下载快照备份 。该快照可用于在需要时将集群恢复至升级前的状态。若用户在升级过程中遇到问题,应首先识别并解决根本原因。若集群仍处于混合版本状态(即至少有一个成员仍运行在 v3.6 版本),可选择将二进制文件或镜像替换为旧版 v3.6 版本,或直接使用快照恢复集群。在此混合状态下,集群仍以 v3.6 版本运行,支持回滚而无需执行正式的降级流程。

然而,一旦所有成员均升级至 v3.7 版本,集群即被视为已完全升级,此时使用二进制文件回滚将不再可行。在此情况下,唯一的恢复选项为从升级前的快照进行恢复,或在升级失败时遵循官方 降级指南 。

升级流程

本示例演示如何升级在本地主机上运行的 3 个成员的 v3.6 etcd 集群。以下输出来自在单个主机上使用三个环回端口对 etcd v3.6.12 和 etcd v3.7.0-rc.0 的实际运行结果。

步骤 1: 检查升级要求

集群是否健康且运行 v3.6.11 或更高版本?

etcdctl --endpoints=localhost:2379,localhost:22379,localhost:32379 endpoint health
<<COMMENT
localhost:2379 is healthy: successfully committed proposal: took = 7.681459ms
localhost:22379 is healthy: successfully committed proposal: took = 7.691750ms
localhost:32379 is healthy: successfully committed proposal: took = 7.698000ms
COMMENT

curl http://localhost:2379/version
<<COMMENT
{"etcdserver":"3.6.12","etcdcluster":"3.6.0","storage":"3.6.0"}
COMMENT

Step 2: 从领导者下载快照备份

下载快照备份 ,以便在出现任何问题时提供回退路径。

etcd 领导者保证拥有最新的应用数据,因此应从领导者获取快照:

for p in 2379 22379 32379; do
  echo -n "localhost:$p leader="
  curl -sL http://localhost:$p/metrics | grep "^etcd_server_is_leader " | awk '{print $2}'
done
<<COMMENT
localhost:2379 leader=1
localhost:22379 leader=0
localhost:32379 leader=0
COMMENT

etcdctl --endpoints=localhost:2379 snapshot save backup.db
<<COMMENT
{"level":"info","ts":"2026-06-02T07:01:41.863225+0300","caller":"snapshot/v3_snapshot.go:83","msg":"created temporary db file","path":"backup.db.part"}
{"level":"info","ts":"2026-06-02T07:01:41.866451+0300","logger":"client","caller":"v3@v3.6.12/maintenance.go:236","msg":"opened snapshot stream; downloading"}
{"level":"info","ts":"2026-06-02T07:01:41.874080+0300","caller":"snapshot/v3_snapshot.go:96","msg":"fetching snapshot","endpoint":"localhost:2379"}
{"level":"info","ts":"2026-06-02T07:01:41.877203+0300","caller":"snapshot/v3_snapshot.go:111","msg":"fetched snapshot","endpoint":"localhost:2379","size":"98 kB","took":"13.822583ms","etcd-version":"3.6.0"}
{"level":"info","ts":"2026-06-02T07:01:41.877303+0300","caller":"snapshot/v3_snapshot.go:121","msg":"saved","path":"backup.db"}
Snapshot saved at backup.db
Server version 3.6.0
COMMENT

第 3 步:停止一个现有的 etcd 服务器

当每个 etcd 进程停止时,集群中的其他成员会记录预期的错误日志。这是正常的,因为集群成员之间的连接已(暂时)中断。领导者将在退出前转移领导权:

{"level":"info","ts":"2026-06-02T07:01:54.949299+0300","caller":"etcdserver/server.go:1274","msg":"leadership transfer finished","local-member-id":"7339c4e5e833c029","old-leader-member-id":"7339c4e5e833c029","new-leader-member-id":"b548c2511513015","took":"101.052625ms"}
{"level":"info","ts":"2026-06-02T07:01:54.949369+0300","caller":"etcdserver/server.go:2349","msg":"server has stopped; stopping cluster version's monitor"}
{"level":"info","ts":"2026-06-02T07:01:55.503219+0300","caller":"embed/etcd.go:626","msg":"stopped serving peer traffic","address":"127.0.0.1:2380"}

第 4 步:使用相同配置重启 etcd 服务器

使用相同配置但采用新 etcd 二进制文件重启 etcd 服务器。

-etcd-old --name ${name} \
+etcd-new --name ${name} \
  --data-dir /path/to/${name}.etcd \
  --listen-client-urls http://localhost:2379 \
  --advertise-client-urls http://localhost:2379 \
  --listen-peer-urls http://localhost:2380 \
  --initial-advertise-peer-urls http://localhost:2380 \
  --initial-cluster s1=http://localhost:2380,s2=http://localhost:22380,s3=http://localhost:32380 \
  --initial-cluster-token tkn \
  --initial-cluster-state new

新的 v3.7 etcd 将向集群发布其信息。此时,集群仍以 v3.6 协议运行,该版本为最低公共版本。

{"level":"info","ts":"2026-06-02T07:01:58.920780+0300","caller":"membership/cluster.go:296","msg":"set cluster version from store","cluster-version":"3.6"}

{"level":"info","ts":"2026-06-02T07:01:58.979186+0300","caller":"etcdserver/server.go:1828","msg":"published local member to cluster through raft","local-member-id":"7339c4e5e833c029","local-member-attributes":"{Name:s1 ClientURLs:[http://localhost:2379]}","cluster-id":"7dee9ba76d59ed53","publish-timeout":"7s"}

验证每个成员以及整个集群在使用新的 v3.7 etcd 二进制文件后是否恢复正常健康状态:

etcdctl --endpoints=localhost:2379,localhost:22379,localhost:32379 endpoint status -w table
<<COMMENT
+-----------------+------------------+------------+-----------------+---------+--------+-----------+
|    ENDPOINT     |        ID        |  VERSION   | STORAGE VERSION | DB SIZE | LEADER | RAFT TERM |
+-----------------+------------------+------------+-----------------+---------+--------+-----------+
|  localhost:2379 | 7339c4e5e833c029 | 3.7.0-rc.0 |           3.6.0 |   98 kB |  false |         3 |
| localhost:22379 | 729934363faa4a24 |     3.6.12 |           3.6.0 |   98 kB |  false |         3 |
| localhost:32379 |  b548c2511513015 |     3.6.12 |           3.6.0 |   98 kB |   true |         3 |
+-----------------+------------------+------------+-----------------+---------+--------+-----------+
COMMENT

未升级的成员和已升级的成员将持续记录关于混合版本状态的日志,直到整个集群完成升级。这是预期行为,当所有 etcd 集群成员均升级至 v3.7 后,日志将停止。

第 5 步:重复第 3 步和第 4 步,对剩余的成员进行操作

所有成员升级完成后,集群将成功报告升级至 v3.7:

{"level":"info","ts":"2026-06-02T07:02:36.054783+0300","caller":"etcdserver/server.go:2311","msg":"updating cluster version using v3 API","from":"3.6","to":"3.7"}

{"level":"info","ts":"2026-06-02T07:02:36.059345+0300","caller":"membership/cluster.go:593","msg":"updated cluster version","cluster-id":"7dee9ba76d59ed53","local-member-id":"7339c4e5e833c029","from":"3.6","to":"3.7"}

{"level":"info","ts":"2026-06-02T07:02:36.059409+0300","caller":"etcdserver/server.go:2326","msg":"cluster version is updated","cluster-version":"3.7"}

etcdctl --endpoints=localhost:2379,localhost:22379,localhost:32379 endpoint health
<<COMMENT
localhost:2379 is healthy: successfully committed proposal: took = 550.833µs
localhost:32379 is healthy: successfully committed proposal: took = 733.458µs
localhost:22379 is healthy: successfully committed proposal: took = 714.416µs
COMMENT

curl http://localhost:2379/version
<<COMMENT
{"etcdserver":"3.7.0-rc.0","etcdcluster":"3.7.0","storage":"3.7.0"}
COMMENT

curl http://localhost:22379/version
<<COMMENT
{"etcdserver":"3.7.0-rc.0","etcdcluster":"3.7.0","storage":"3.7.0"}
COMMENT

curl http://localhost:32379/version
<<COMMENT
{"etcdserver":"3.7.0-rc.0","etcdcluster":"3.7.0","storage":"3.7.0"}
COMMENT

6 - 将 etcd 从 3.2 升级到 3.3

升级 etcd 3.2 至 3.3 的流程、检查清单与注意事项

在一般情况下,从 etcd 3.2 升级到 3.3 可以实现零停机滚动升级:

  • 逐一停止 etcd v3.2 进程,并替换为 etcd v3.3 进程
  • 在所有 v3.3 进程运行后,集群即可使用 v3.3 的新特性

在 开始升级 之前,请通读本指南其余部分以做好准备。

升级检查列表

警告

从 没有 v3 数据的 v2 迁移 时,如果 etcd 从现有快照恢复,但不存在 v3 ETCD_DATA_DIR/member/snap/db 文件,etcd v3.2+ 服务器会发生崩溃。这种情况出现在服务器由 v2 迁移且此前没有 v3 数据时。此限制也可防止意外丢失 v3 数据(例如 db 文件可能已被移动)。etcd 要求 v3 迁移后的操作必须有 v3 数据。v3.0 服务器包含 v3 数据之前,请勿升级到更新的 v3 版本。

警告

若启用认证并使用租约(租约 TTL 较小),极有可能遇到 问题 ,导致数据不一致。强烈建议先升级至 3.2.31+ 以修复此问题,再升级至 3.3。此外,在升级过程中,若无权限用户向 3.3 节点发送 LeaseRevoke 请求,仍可能导致数据损坏,因此建议在升级前确保环境中不存在此类异常调用,详情请参见 #11691 。

3.3 版本中的重点变更。

将 etcd --auto-compaction-retention 标志的价值类型更改为 string

将 --auto-compaction-retention 标志改为 接受字符串值 ,并支持 更细粒度 。由于 --auto-compaction-retention 现在接受字符串值,etcd 配置 YAML 文件 auto-compaction-retention 字段必须改为 string 类型。此前 --config-file etcd.config.yaml 可包含 auto-compaction-retention: 24 字段,现在必须为 auto-compaction-retention: "24" 或 auto-compaction-retention: "24h"。若配置为 --auto-compaction-mode periodic --auto-compaction-retention "24h",则 --auto-compaction-retention 标志的时间持续值必须对 Go 中的 time.ParseDuration 函数有效。

# etcd.config.yaml
+auto-compaction-mode: periodic
-auto-compaction-retention: 24
+auto-compaction-retention: "24"
+# Or
+auto-compaction-retention: "24h"

将 etcdserver.EtcdServer.ServerConfig 更改为 *etcdserver.EtcdServer.ServerConfig

etcdserver.EtcdServer 已将成员字段 *etcdserver.ServerConfig 的类型更改为 etcdserver.ServerConfig。现在 etcdserver.NewServer 接受 etcdserver.ServerConfig,而非 *etcdserver.ServerConfig。

之前和之后(例如 k8s.io/kubernetes/test/e2e_node/services/etcd.go )

import "github.com/coreos/etcd/etcdserver"

type EtcdServer struct {
	*etcdserver.EtcdServer
-	config *etcdserver.ServerConfig
+	config etcdserver.ServerConfig
}

func NewEtcd(dataDir string) *EtcdServer {
-	config := &etcdserver.ServerConfig{
+	config := etcdserver.ServerConfig{
		DataDir: dataDir,
        ...
	}
	return &EtcdServer{config: config}
}

func (e *EtcdServer) Start() error {
	var err error
	e.EtcdServer, err = etcdserver.NewServer(e.config)
    ...

添加了 embed.Config.LogOutput 结构体

警告

请注意,此字段在 v3.4 版本中已重命名为 embed.Config.LogOutputs,适用于 []string 类型。详情请参阅 v3.4 升级指南 。

字段 LogOutput 已添加至 embed.Config:

package embed

type Config struct {
 	Debug bool `json:"debug"`
 	LogPkgLevels string `json:"log-package-levels"`
+	LogOutput string `json:"log-output"`
 	...

在 gRPC 服务器警告被记录到 etcdserver 之前。

WARNING: 2017/11/02 11:35:51 grpc: addrConn.resetTransport failed to create client transport: connection error: desc = "transport: Error while dialing dial tcp: operation was canceled"; Reconnecting to {localhost:2379 <nil>}
WARNING: 2017/11/02 11:35:51 grpc: addrConn.resetTransport failed to create client transport: connection error: desc = "transport: Error while dialing dial tcp: operation was canceled"; Reconnecting to {localhost:2379 <nil>}

从 v3.3 版本开始,gRPC 服务器日志默认已禁用。

警告

请注意,embed.Config.SetupLogging 方法已在 v3.4 版本中弃用。详情请参阅 v3.4 升级指南 。

import "github.com/coreos/etcd/embed"

cfg := &embed.Config{Debug: false}
cfg.SetupLogging()

将 embed.Config.Debug 字段设置为 true 以启用 gRPC 服务器日志。

Changed /health 端点响应

此前,[endpoint]:[client-port]/health 返回手动序列化的 JSON 值。3.3 版本现在定义了 etcdhttp.Health 结构体。

请注意,在 v3.3.0-rc.0、v3.3.0-rc.1 和 v3.3.0-rc.2 版本中,etcdhttp.Health 的 "health" 和 "errors" 字段为布尔类型。为保持向后兼容性,已将 "health" 字段恢复为 string 类型,并移除了 "errors" 字段。后续的健康信息将通过独立的 API 提供。

$ curl http://localhost:2379/health
{"health":"true"}

Changed gRPC 网关 HTTP 端点(替换 /v3alpha 为 /v3beta)

本文未提供内容。

curl -L http://localhost:2379/v3alpha/kv/put \
  -X POST -d '{"key": "Zm9v", "value": "YmFy"}'

之后

curl -L http://localhost:2379/v3beta/kv/put \
  -X POST -d '{"key": "Zm9v", "value": "YmFy"}'

对 /v3alpha 端点的请求将重定向至 /v3beta,/v3alpha 将在 3.4 版本中移除。

调整了最大请求大小限制

3.3 现在允许为服务器端和客户端分别设置自定义请求大小限制。在之前版本(v3.2.10、v3.2.11)中,客户端响应大小限制仅为 4 MiB。

服务器端请求限制可通过 --max-request-bytes 标志进行配置:

# limits request size to 1.5 KiB
etcd --max-request-bytes 1536

# client writes exceeding 1.5 KiB will be rejected
etcdctl put foo [LARGE VALUE...]
# etcdserver: request is too large

或配置 embed.Config.MaxRequestBytes 字段:

import "github.com/coreos/etcd/embed"
import "github.com/coreos/etcd/etcdserver/api/v3rpc/rpctypes"

// limit requests to 5 MiB
cfg := embed.NewConfig()
cfg.MaxRequestBytes = 5 * 1024 * 1024

// client writes exceeding 5 MiB will be rejected
_, err := cli.Put(ctx, "foo", [LARGE VALUE...])
err == rpctypes.ErrRequestTooLarge

如果未指定,服务器端限制默认为 1.5 MiB。

客户端请求限制必须根据服务器端限制进行配置。

# limits request size to 1 MiB
etcd --max-request-bytes 1048576
import "github.com/coreos/etcd/clientv3"

cli, _ := clientv3.New(clientv3.Config{
    Endpoints: []string{"127.0.0.1:2379"},
    MaxCallSendMsgSize: 2 * 1024 * 1024,
    MaxCallRecvMsgSize: 3 * 1024 * 1024,
})


// client writes exceeding "--max-request-bytes" will be rejected from etcd server
_, err := cli.Put(ctx, "foo", strings.Repeat("a", 1*1024*1024+5))
err == rpctypes.ErrRequestTooLarge


// client writes exceeding "MaxCallSendMsgSize" will be rejected from client-side
_, err = cli.Put(ctx, "foo", strings.Repeat("a", 5*1024*1024))
err.Error() == "rpc error: code = ResourceExhausted desc = grpc: trying to send message larger than max (5242890 vs. 2097152)"


// some writes under limits
for i := range []int{0,1,2,3,4} {
    _, err = cli.Put(ctx, fmt.Sprintf("foo%d", i), strings.Repeat("a", 1*1024*1024-500))
    if err != nil {
        panic(err)
    }
}
// client reads exceeding "MaxCallRecvMsgSize" will be rejected from client-side
_, err = cli.Get(ctx, "foo", clientv3.WithPrefix())
err.Error() == "rpc error: code = ResourceExhausted desc = grpc: received message larger than max (5240509 vs. 3145728)"

如果未指定,客户端发送限制默认为 2 MiB(1.5 MiB + gRPC 开销字节),接收限制为 math.MaxInt32。请参阅 clientv3 godoc 获取更多详细信息。

更改了原始 gRPC 客户端包装函数的签名

3.3 修改了 clientv3 gRPC 客户端封装的函数签名。此变更旨在支持 自定义 grpc.CallOption 消息大小限制 。

之前和之后

-func NewKVFromKVClient(remote pb.KVClient) KV {
+func NewKVFromKVClient(remote pb.KVClient, c *Client) KV {

-func NewClusterFromClusterClient(remote pb.ClusterClient) Cluster {
+func NewClusterFromClusterClient(remote pb.ClusterClient, c *Client) Cluster {

-func NewLeaseFromLeaseClient(remote pb.LeaseClient, keepAliveTimeout time.Duration) Lease {
+func NewLeaseFromLeaseClient(remote pb.LeaseClient, c *Client, keepAliveTimeout time.Duration) Lease {

-func NewMaintenanceFromMaintenanceClient(remote pb.MaintenanceClient) Maintenance {
+func NewMaintenanceFromMaintenanceClient(remote pb.MaintenanceClient, c *Client) Maintenance {

-func NewWatchFromWatchClient(wc pb.WatchClient) Watcher {
+func NewWatchFromWatchClient(wc pb.WatchClient, c *Client) Watcher {

Changed clientv3 Snapshot API 错误类型

此前,clientv3 Snapshot API 返回原始的 [grpc/*status.statusError] 类型错误。v3.3 现已将这些错误转换为对应的公开错误类型,以与其他 API 保持一致。

本文未提供内容。

import "context"

// reading snapshot with canceled context should error out
ctx, cancel := context.WithCancel(context.Background())
rc, _ := cli.Snapshot(ctx)
cancel()
_, err := io.Copy(f, rc)
err.Error() == "rpc error: code = Canceled desc = context canceled"

// reading snapshot with deadline exceeded should error out
ctx, cancel = context.WithTimeout(context.Background(), time.Second)
defer cancel()
rc, _ = cli.Snapshot(ctx)
time.Sleep(2 * time.Second)
_, err = io.Copy(f, rc)
err.Error() == "rpc error: code = DeadlineExceeded desc = context deadline exceeded"

之后

import "context"

// reading snapshot with canceled context should error out
ctx, cancel := context.WithCancel(context.Background())
rc, _ := cli.Snapshot(ctx)
cancel()
_, err := io.Copy(f, rc)
err == context.Canceled

// reading snapshot with deadline exceeded should error out
ctx, cancel = context.WithTimeout(context.Background(), time.Second)
defer cancel()
rc, _ = cli.Snapshot(ctx)
time.Sleep(2 * time.Second)
_, err = io.Copy(f, rc)
err == context.DeadlineExceeded

Changed etcdctl lease timetolive 命令输出

此前,对已过期租约执行 lease timetolive LEASE_ID 命令时会输出 -1s 表示剩余秒数。3.3 版本现在输出更清晰的提示信息。

本文未提供内容。

lease 2d8257079fa1bc0c granted with TTL(0s), remaining(-1s)

之后

lease 2d8257079fa1bc0c already expired

变更 golang.org/x/net/context 导入

clientv3 已弃用 golang.org/x/net/context。若项目在其他代码中引入 golang.org/x/net/context(例如 etcd 生成的协议缓冲区代码)并导入 github.com/coreos/etcd/clientv3,则编译时需使用 Go 1.9 或更高版本。

本文未提供内容。

import "golang.org/x/net/context"
cli.Put(context.Background(), "f", "v")

之后

import "context"
cli.Put(context.Background(), "f", "v")

更改了 gRPC 依赖

3.3 必须使用 grpc/grpc-go v1.7.5。

已弃用 grpclog.Logger

grpclog.Logger 已被弃用,建议改用 grpclog.LoggerV2 。clientv3.Logger 现已改为 grpclog.LoggerV2。

本文未提供内容。

import "github.com/coreos/etcd/clientv3"
clientv3.SetLogger(log.New(os.Stderr, "grpc: ", 0))

之后

import "github.com/coreos/etcd/clientv3"
import "google.golang.org/grpc/grpclog"
clientv3.SetLogger(grpclog.NewLoggerV2(os.Stderr, os.Stderr, os.Stderr))

// log.New above cannot be used (not implement grpclog.LoggerV2 interface)
已弃用 grpc.ErrClientConnTimeout

此前,在客户端连接超时时返回 grpc.ErrClientConnTimeout 错误。3.3 版本改为返回 context.DeadlineExceeded(参见 #8504 )。

本文未提供内容。

// expect dial time-out on ipv4 blackhole
_, err := clientv3.New(clientv3.Config{
    Endpoints:   []string{"http://254.0.0.1:12345"},
    DialTimeout: 2 * time.Second
})
if err == grpc.ErrClientConnTimeout {
	// handle errors
}

之后

_, err := clientv3.New(clientv3.Config{
    Endpoints:   []string{"http://254.0.0.1:12345"},
    DialTimeout: 2 * time.Second
})
if err == context.DeadlineExceeded {
	// handle errors
}

变更官方容器注册表

etcd 现在使用 gcr.io/etcd-development/etcd 作为主容器注册表,使用 quay.io/coreos/etcd 作为备用。

本文未提供内容。

docker pull quay.io/coreos/etcd:v3.2.5

之后

docker pull gcr.io/etcd-development/etcd:v3.3.0

升级至>=v3.3.14

v3.3.14 在尽量减少客户端负载均衡实现差异的前提下,引入了部分 3.4 版本的特性。此版本修复了 “当首个 etcd-server 不可用时,kube-apiserver 1.13.x 拒绝运行”(kubernetes#72102) 的问题。

grpc.ErrClientConnClosing 在 gRPC ≥ 1.10 中已被 弃用 。

import (
+	"go.etcd.io/etcd/clientv3"

	"google.golang.org/grpc"
+	"google.golang.org/grpc/codes"
+	"google.golang.org/grpc/status"
)

_, err := kvc.Get(ctx, "a")
-if err == grpc.ErrClientConnClosing {
+if clientv3.IsConnCanceled(err) {

// or
+s, ok := status.FromError(err)
+if ok {
+  if s.Code() == codes.Canceled

新的客户端负载均衡器 使用异步解析器将端点传递给 gRPC dial 函数。因此,v3.3.14 或更高版本必须使用 grpc.WithBlock dial 选项,以等待底层连接建立完成。

import (
	"time"
	"go.etcd.io/etcd/clientv3"
+	"google.golang.org/grpc"
)

+// "grpc.WithBlock()" to block until the underlying connection is up
ccfg := clientv3.Config{
  Endpoints:            []string{"localhost:2379"},
  DialTimeout:          time.Second,
+ DialOptions:          []grpc.DialOption{grpc.WithBlock()},
  DialKeepAliveTime:    time.Second,
  DialKeepAliveTimeout: 500 * time.Millisecond,
}

请参阅 CHANGELOG 以获取完整变更列表。

服务器升级检查清单

升级要求

要将现有 etcd 部署升级至 3.3 版本,运行中的集群版本必须为 3.2 或更高。若版本低于 3.2,请先 升级至 3.2 ,再升级至 3.3。

此外,为确保滚动升级顺利进行,运行中的集群必须处于健康状态。在继续操作前,请使用 etcdctl endpoint health 命令检查集群健康状况。

准备

在升级 etcd 之前,请务必在预发环境中测试依赖 etcd 的服务,再将升级部署到生产环境。

升级前,请对 etcd 数据执行 备份 etcd 数据 。若升级过程中出现异常,可使用此备份将系统 降级 至现有 etcd 版本。请注意,snapshot命令仅备份 v3 数据。如需备份 v2 数据,请参阅 备份 v2 数据存储 。

混合版本

升级期间,etcd 集群支持不同版本的 etcd 成员共存,并以最低公共版本的协议运行。只有当集群中所有成员均升级至 3.3 版本后,该集群才被视为已完成升级。内部机制上,etcd 成员之间会相互协商以确定集群的整体版本,该版本控制报告的版本及所支持的功能。

限制

请注意:如果集群仅包含 v3 数据且无 v2 数据,则不受此限制影响。

如果集群正在服务的数据集大小超过 50MB,每个新升级的成员可能需要最多 2 分钟才能追上现有集群。请检查最近快照的大小以估算总数据量。换句话说,升级每个成员之间应至少等待 2 分钟。

对于数据总量更大(例如 100MB 或更多)的情况,此一次性操作可能需要更长时间。对于规模达到此类程度的大型 etcd 集群,系统管理员可在升级前自由联系 etcd 团队 ,我们将乐意提供升级流程方面的建议。

降级

如果所有成员均已升级至 v3.3 版本,集群将升级至 v3.3 版本,从该完成状态回退不可行。然而,若任一成员仍为 v3.2 版本,则集群及其操作仍保持 “v3.2” 状态,此时可从该混合集群状态恢复至所有成员均使用 v3.2 etcd 二进制文件。

请备份所有 etcd 成员的数据目录 backup the data directory ,以确保在集群完全升级后仍可执行降级操作。

升级流程

本示例演示如何升级在本地计算机上运行的 3 个成员的 v3.2 etcd 集群。

1. 检查升级要求

集群是否健康且运行 v3.2.x 版本?

$ ETCDCTL_API=3 etcdctl endpoint health --endpoints=localhost:2379,localhost:22379,localhost:32379
localhost:2379 is healthy: successfully committed proposal: took = 6.600684ms
localhost:22379 is healthy: successfully committed proposal: took = 8.540064ms
localhost:32379 is healthy: successfully committed proposal: took = 8.763432ms

$ curl http://localhost:2379/version
{"etcdserver":"3.2.7","etcdcluster":"3.2.0"}

2. 停止现有 etcd 进程

当每个 etcd 进程停止时,集群中的其他成员会记录预期的错误。这是正常的,因为集群成员之间的连接已(暂时)中断:

14:13:31.491746 I | raft: c89feb932daef420 [term 3] received MsgTimeoutNow from 6d4f535bae3ab960 and starts an election to get leadership.
14:13:31.491769 I | raft: c89feb932daef420 became candidate at term 4
14:13:31.491788 I | raft: c89feb932daef420 received MsgVoteResp from c89feb932daef420 at term 4
14:13:31.491797 I | raft: c89feb932daef420 [logterm: 3, index: 9] sent MsgVote request to 6d4f535bae3ab960 at term 4
14:13:31.491805 I | raft: c89feb932daef420 [logterm: 3, index: 9] sent MsgVote request to 9eda174c7df8a033 at term 4
14:13:31.491815 I | raft: raft.node: c89feb932daef420 lost leader 6d4f535bae3ab960 at term 4
14:13:31.524084 I | raft: c89feb932daef420 received MsgVoteResp from 6d4f535bae3ab960 at term 4
14:13:31.524108 I | raft: c89feb932daef420 [quorum:2] has received 2 MsgVoteResp votes and 0 vote rejections
14:13:31.524123 I | raft: c89feb932daef420 became leader at term 4
14:13:31.524136 I | raft: raft.node: c89feb932daef420 elected leader c89feb932daef420 at term 4
14:13:31.592650 W | rafthttp: lost the TCP streaming connection with peer 6d4f535bae3ab960 (stream MsgApp v2 reader)
14:13:31.592825 W | rafthttp: lost the TCP streaming connection with peer 6d4f535bae3ab960 (stream Message reader)
14:13:31.693275 E | rafthttp: failed to dial 6d4f535bae3ab960 on stream Message (dial tcp [::1]:2380: getsockopt: connection refused)
14:13:31.693289 I | rafthttp: peer 6d4f535bae3ab960 became inactive
14:13:31.936678 W | rafthttp: lost the TCP streaming connection with peer 6d4f535bae3ab960 (stream Message writer)

此时建议 备份 etcd 数据 ,以便在出现任何问题时可回退至之前版本:

$ etcdctl snapshot save backup.db

3. 插入 etcd v3.3 二进制文件并启动新 etcd 进程

新的 v3.3 版 etcd 将向集群发布其信息:

14:14:25.363225 I | etcdserver: published {Name:s1 ClientURLs:[http://localhost:2379]} to cluster a9ededbffcb1b1f1

验证每个成员以及整个集群在使用新的 v3.3 etcd 二进制文件后是否恢复正常健康状态:

$ ETCDCTL_API=3 /etcdctl endpoint health --endpoints=localhost:2379,localhost:22379,localhost:32379
localhost:22379 is healthy: successfully committed proposal: took = 5.540129ms
localhost:32379 is healthy: successfully committed proposal: took = 7.321771ms
localhost:2379 is healthy: successfully committed proposal: took = 10.629901ms

升级后的成员将在整个集群完成升级前持续记录类似以下的警告日志。这是预期行为,当所有 etcd 集群成员均升级至 v3.3 后,警告将停止出现。

14:15:17.071804 W | etcdserver: member c89feb932daef420 has a higher version 3.3.0
14:15:21.073110 W | etcdserver: the local etcd version 3.2.7 is not up-to-date
14:15:21.073142 W | etcdserver: member 6d4f535bae3ab960 has a higher version 3.3.0
14:15:21.073157 W | etcdserver: the local etcd version 3.2.7 is not up-to-date
14:15:21.073164 W | etcdserver: member c89feb932daef420 has a higher version 3.3.0

4. 重复第 2 步到第 3 步,对所有其他成员执行

5. 完成

所有成员升级完成后,集群将成功报告升级至 3.3:

14:15:54.536901 N | etcdserver/membership: updated the cluster version from 3.2 to 3.3
14:15:54.537035 I | etcdserver/api: enabled capabilities for version 3.3
$ ETCDCTL_API=3 /etcdctl endpoint health --endpoints=localhost:2379,localhost:22379,localhost:32379
localhost:2379 is healthy: successfully committed proposal: took = 2.312897ms
localhost:22379 is healthy: successfully committed proposal: took = 2.553476ms
localhost:32379 is healthy: successfully committed proposal: took = 2.517902ms

7 - 将 etcd 从 3.1 升级到 3.2

升级 etcd 3.1 至 3.2 的流程、检查清单与注意事项

在一般情况下,从 etcd 3.1 升级到 3.2 可以实现零停机滚动升级:

  • 逐一停止 etcd v3.1 进程,并替换为 etcd v3.2 进程
  • 在所有 v3.2 进程运行后,集群即可使用 v3.2 的新特性

在 开始升级 之前,请通读本指南其余部分以做好准备。

升级检查列表

警告

从 没有 v3 数据的 v2 迁移 时,如果 etcd 从现有快照恢复,但不存在 v3 ETCD_DATA_DIR/member/snap/db 文件,etcd v3.2+ 服务器会发生崩溃。这种情况出现在服务器由 v2 迁移且此前没有 v3 数据时。此限制也可防止意外丢失 v3 数据(例如 db 文件可能已被移动)。etcd 要求 v3 迁移后的操作必须有 v3 数据。v3.0 服务器包含 v3 数据之前,请勿升级到更新的 v3 版本。

3.2 版本中的重点变更。

Changed default snapshot-count value

较高的 --snapshot-count 会在生成快照前将更多 Raft 条目保留在内存中,从而导致 内存使用量持续较高 。由于领导者会更长时间保留最新的 Raft 条目,缓慢的跟随者有更多时间在领导者生成快照前完成追赶。--snapshot-count 是较高内存使用与缓慢跟随者更高可用性之间的权衡。

自 v3.2 起,--snapshot-count 的默认值已 从 10,000 改为 100,000 。

更新了 gRPC 依赖 (>=3.2.10)

3.2.10 及更高版本现在要求 grpc/grpc-go v1.7.5(3.2.9 及更早版本要求 v1.2.1)。

已弃用 grpclog.Logger

grpclog.Logger 已被弃用,建议改用 grpclog.LoggerV2 。clientv3.Logger 现已改为 grpclog.LoggerV2。

本文未提供内容。

import "github.com/coreos/etcd/clientv3"
clientv3.SetLogger(log.New(os.Stderr, "grpc: ", 0))

之后

import "github.com/coreos/etcd/clientv3"
import "google.golang.org/grpc/grpclog"
clientv3.SetLogger(grpclog.NewLoggerV2(os.Stderr, os.Stderr, os.Stderr))

// log.New above cannot be used (not implement grpclog.LoggerV2 interface)
已弃用 grpc.ErrClientConnTimeout

此前,在客户端连接超时时返回 grpc.ErrClientConnTimeout 错误。3.2 版本改为返回 context.DeadlineExceeded(参见 #8504 )。

本文未提供内容。

// expect dial time-out on ipv4 blackhole
_, err := clientv3.New(clientv3.Config{
    Endpoints:   []string{"http://254.0.0.1:12345"},
    DialTimeout: 2 * time.Second
})
if err == grpc.ErrClientConnTimeout {
	// handle errors
}

之后

_, err := clientv3.New(clientv3.Config{
    Endpoints:   []string{"http://254.0.0.1:12345"},
    DialTimeout: 2 * time.Second
})
if err == context.DeadlineExceeded {
	// handle errors
}

调整了最大请求大小限制(>=3.2.10)

3.2.10 和 3.2.11 版本允许在服务端自定义请求大小限制。从 3.2.12 版本开始,服务端和客户端均支持自定义请求大小限制。在之前的版本(v3.2.10、v3.2.11)中,客户端响应大小仅限于 4 MiB。

服务器端请求限制可通过 --max-request-bytes 标志进行配置:

# limits request size to 1.5 KiB
etcd --max-request-bytes 1536

# client writes exceeding 1.5 KiB will be rejected
etcdctl put foo [LARGE VALUE...]
# etcdserver: request is too large

或配置 embed.Config.MaxRequestBytes 字段:

import "github.com/coreos/etcd/embed"
import "github.com/coreos/etcd/etcdserver/api/v3rpc/rpctypes"

// limit requests to 5 MiB
cfg := embed.NewConfig()
cfg.MaxRequestBytes = 5 * 1024 * 1024

// client writes exceeding 5 MiB will be rejected
_, err := cli.Put(ctx, "foo", [LARGE VALUE...])
err == rpctypes.ErrRequestTooLarge

如果未指定,服务器端限制默认为 1.5 MiB。

客户端请求限制必须根据服务器端限制进行配置。

# limits request size to 1 MiB
etcd --max-request-bytes 1048576
import "github.com/coreos/etcd/clientv3"

cli, _ := clientv3.New(clientv3.Config{
    Endpoints: []string{"127.0.0.1:2379"},
    MaxCallSendMsgSize: 2 * 1024 * 1024,
    MaxCallRecvMsgSize: 3 * 1024 * 1024,
})


// client writes exceeding "--max-request-bytes" will be rejected from etcd server
_, err := cli.Put(ctx, "foo", strings.Repeat("a", 1*1024*1024+5))
err == rpctypes.ErrRequestTooLarge


// client writes exceeding "MaxCallSendMsgSize" will be rejected from client-side
_, err = cli.Put(ctx, "foo", strings.Repeat("a", 5*1024*1024))
err.Error() == "rpc error: code = ResourceExhausted desc = grpc: trying to send message larger than max (5242890 vs. 2097152)"


// some writes under limits
for i := range []int{0,1,2,3,4} {
    _, err = cli.Put(ctx, fmt.Sprintf("foo%d", i), strings.Repeat("a", 1*1024*1024-500))
    if err != nil {
        panic(err)
    }
}
// client reads exceeding "MaxCallRecvMsgSize" will be rejected from client-side
_, err = cli.Get(ctx, "foo", clientv3.WithPrefix())
err.Error() == "rpc error: code = ResourceExhausted desc = grpc: received message larger than max (5240509 vs. 3145728)"

如果未指定,客户端发送限制默认为 2 MiB(1.5 MiB + gRPC 开销字节),接收限制为 math.MaxInt32。请参阅 clientv3 godoc 获取更多详细信息。

更改了原始 gRPC 客户端包装器

3.2.12 及更高版本更改了 clientv3 gRPC 客户端封装的函数签名。此更改旨在支持 自定义 grpc.CallOption 消息大小限制 。

之前和之后

-func NewKVFromKVClient(remote pb.KVClient) KV {
+func NewKVFromKVClient(remote pb.KVClient, c *Client) KV {

-func NewClusterFromClusterClient(remote pb.ClusterClient) Cluster {
+func NewClusterFromClusterClient(remote pb.ClusterClient, c *Client) Cluster {

-func NewLeaseFromLeaseClient(remote pb.LeaseClient, keepAliveTimeout time.Duration) Lease {
+func NewLeaseFromLeaseClient(remote pb.LeaseClient, c *Client, keepAliveTimeout time.Duration) Lease {

-func NewMaintenanceFromMaintenanceClient(remote pb.MaintenanceClient) Maintenance {
+func NewMaintenanceFromMaintenanceClient(remote pb.MaintenanceClient, c *Client) Maintenance {

-func NewWatchFromWatchClient(wc pb.WatchClient) Watcher {
+func NewWatchFromWatchClient(wc pb.WatchClient, c *Client) Watcher {

变更 clientv3.Lease.TimeToLive API

此前,clientv3.Lease.TimeToLive API 在不存在的租约 ID 上返回 lease.ErrLeaseNotFound。3.2 版本改为在响应中返回 TTL=-1 且不返回错误(参见 #7305 )。

本文未提供内容。

// when leaseID does not exist
resp, err := TimeToLive(ctx, leaseID)
resp == nil
err == lease.ErrLeaseNotFound

之后

// when leaseID does not exist
resp, err := TimeToLive(ctx, leaseID)
resp.TTL == -1
err == nil

将clientv3.NewFromConfigFile移动到clientv3.yaml.NewConfig

clientv3.NewFromConfigFile 已移至 yaml.NewConfig。

本文未提供内容。

import "github.com/coreos/etcd/clientv3"
clientv3.NewFromConfigFile

之后

import clientv3yaml "github.com/coreos/etcd/clientv3/yaml"
clientv3yaml.NewConfig

Change in --listen-peer-urls and --listen-client-urls

3.2 现在拒绝为 --listen-peer-urls 和 --listen-client-urls 使用域名(3.1 仅输出警告),因为域名对网络接口绑定无效。请确保这些 URL 已正确格式化为 scheme://IP:port。

有关更多上下文,请参见 issue #6336 。

服务器升级检查清单

升级要求

要将现有 etcd 部署升级至 3.2 版本,运行中的集群版本必须为 3.1 或更高。若版本低于 3.1,请先 升级至 3.1 ,再升级至 3.2。

此外,为确保滚动升级顺利进行,运行中的集群必须处于健康状态。在继续操作前,请使用 etcdctl endpoint health 命令检查集群健康状况。

准备

在升级 etcd 之前,请务必在预发环境中测试依赖 etcd 的服务,再将升级部署到生产环境。

升级前,请对 etcd 数据执行 备份 etcd 数据 。若升级过程中出现异常,可使用此备份将系统 降级 至现有 etcd 版本。请注意,snapshot命令仅备份 v3 数据。如需备份 v2 数据,请参阅 备份 v2 数据存储 。

混合版本

升级期间,etcd 集群支持不同版本的 etcd 成员共存,并以最低公共版本的协议运行。只有当集群中所有成员均升级至 3.2 版本后,才认为集群已完成升级。内部机制上,etcd 成员之间会相互协商以确定集群的整体版本,该版本控制报告的版本及支持的功能。

限制

请注意:如果集群仅包含 v3 数据且无 v2 数据,则不受此限制影响。

如果集群正在服务的数据集大小超过 50MB,每个新升级的成员可能需要最多 2 分钟才能追上现有集群。请检查最近快照的大小以估算总数据量。换句话说,升级每个成员之间应至少等待 2 分钟。

对于数据总量更大(例如 100MB 或更多)的情况,此一次性操作可能需要更长时间。对于规模达到此类程度的大型 etcd 集群,系统管理员可在升级前自由联系 etcd 团队 ,我们将乐意提供升级流程方面的建议。

降级

如果所有成员均已升级至 v3.2 版本,集群将升级至 v3.2 版本,从该完成状态回退不可行。然而,若任一成员仍为 v3.1 版本,则集群及其操作仍保持 “v3.1” 状态,此时可从该混合集群状态恢复至所有成员均使用 v3.1 etcd 二进制文件。

请注意,务必对所有 etcd 成员的数据目录 backup the data directory 进行备份,以确保在集群完全升级后仍可执行降级操作。

升级流程

本示例演示如何升级在本地计算机上运行的 3 个成员的 v3.1 etcd 集群。

1. 检查升级要求

集群是否健康且运行 v3.1.x 版本?

$ ETCDCTL_API=3 etcdctl endpoint health --endpoints=localhost:2379,localhost:22379,localhost:32379
localhost:2379 is healthy: successfully committed proposal: took = 6.600684ms
localhost:22379 is healthy: successfully committed proposal: took = 8.540064ms
localhost:32379 is healthy: successfully committed proposal: took = 8.763432ms

$ curl http://localhost:2379/version
{"etcdserver":"3.1.7","etcdcluster":"3.1.0"}

2. 停止现有 etcd 进程

当每个 etcd 进程停止时,集群中的其他成员会记录预期的错误。这是正常的,因为集群成员之间的连接已(暂时)中断:

2017-04-27 14:13:31.491746 I | raft: c89feb932daef420 [term 3] received MsgTimeoutNow from 6d4f535bae3ab960 and starts an election to get leadership.
2017-04-27 14:13:31.491769 I | raft: c89feb932daef420 became candidate at term 4
2017-04-27 14:13:31.491788 I | raft: c89feb932daef420 received MsgVoteResp from c89feb932daef420 at term 4
2017-04-27 14:13:31.491797 I | raft: c89feb932daef420 [logterm: 3, index: 9] sent MsgVote request to 6d4f535bae3ab960 at term 4
2017-04-27 14:13:31.491805 I | raft: c89feb932daef420 [logterm: 3, index: 9] sent MsgVote request to 9eda174c7df8a033 at term 4
2017-04-27 14:13:31.491815 I | raft: raft.node: c89feb932daef420 lost leader 6d4f535bae3ab960 at term 4
2017-04-27 14:13:31.524084 I | raft: c89feb932daef420 received MsgVoteResp from 6d4f535bae3ab960 at term 4
2017-04-27 14:13:31.524108 I | raft: c89feb932daef420 [quorum:2] has received 2 MsgVoteResp votes and 0 vote rejections
2017-04-27 14:13:31.524123 I | raft: c89feb932daef420 became leader at term 4
2017-04-27 14:13:31.524136 I | raft: raft.node: c89feb932daef420 elected leader c89feb932daef420 at term 4
2017-04-27 14:13:31.592650 W | rafthttp: lost the TCP streaming connection with peer 6d4f535bae3ab960 (stream MsgApp v2 reader)
2017-04-27 14:13:31.592825 W | rafthttp: lost the TCP streaming connection with peer 6d4f535bae3ab960 (stream Message reader)
2017-04-27 14:13:31.693275 E | rafthttp: failed to dial 6d4f535bae3ab960 on stream Message (dial tcp [::1]:2380: getsockopt: connection refused)
2017-04-27 14:13:31.693289 I | rafthttp: peer 6d4f535bae3ab960 became inactive
2017-04-27 14:13:31.936678 W | rafthttp: lost the TCP streaming connection with peer 6d4f535bae3ab960 (stream Message writer)

此时建议 备份 etcd 数据 ,以便在出现任何问题时提供回退路径:

$ etcdctl snapshot save backup.db

3. 直接部署 etcd v3.2 二进制文件并启动新 etcd 进程

新的 v3.2 版 etcd 将向集群发布其信息:

2017-04-27 14:14:25.363225 I | etcdserver: published {Name:s1 ClientURLs:[http://localhost:2379]} to cluster a9ededbffcb1b1f1

验证每个成员,然后整个集群,使用新的 v3.2 etcd 二进制文件后是否健康:

$ ETCDCTL_API=3 /etcdctl endpoint health --endpoints=localhost:2379,localhost:22379,localhost:32379
localhost:22379 is healthy: successfully committed proposal: took = 5.540129ms
localhost:32379 is healthy: successfully committed proposal: took = 7.321771ms
localhost:2379 is healthy: successfully committed proposal: took = 10.629901ms

升级后的成员将在整个集群完成升级前持续记录如下警告信息。这是正常现象,待所有 etcd 集群成员均升级至 v3.2 后,警告将停止出现。

2017-04-27 14:15:17.071804 W | etcdserver: member c89feb932daef420 has a higher version 3.2.0
2017-04-27 14:15:21.073110 W | etcdserver: the local etcd version 3.1.7 is not up-to-date
2017-04-27 14:15:21.073142 W | etcdserver: member 6d4f535bae3ab960 has a higher version 3.2.0
2017-04-27 14:15:21.073157 W | etcdserver: the local etcd version 3.1.7 is not up-to-date
2017-04-27 14:15:21.073164 W | etcdserver: member c89feb932daef420 has a higher version 3.2.0

4. 重复第 2 步到第 3 步,对所有其他成员执行

5. 完成

所有成员升级完成后,集群将成功报告升级至 3.2:

2017-04-27 14:15:54.536901 N | etcdserver/membership: updated the cluster version from 3.1 to 3.2
2017-04-27 14:15:54.537035 I | etcdserver/api: enabled capabilities for version 3.2
$ ETCDCTL_API=3 /etcdctl endpoint health --endpoints=localhost:2379,localhost:22379,localhost:32379
localhost:2379 is healthy: successfully committed proposal: took = 2.312897ms
localhost:22379 is healthy: successfully committed proposal: took = 2.553476ms
localhost:32379 is healthy: successfully committed proposal: took = 2.517902ms

8 - 将 etcd 从 3.0 升级到 3.1

升级 etcd 3.0 至 3.1 的流程、检查清单与注意事项

在一般情况下,从 etcd 3.0 升级到 3.1 可以实现零停机滚动升级:

  • 逐一停止 etcd v3.0 进程,并替换为 etcd v3.1 进程
  • 在所有 v3.1 进程运行后,集群即可使用 v3.1 的新特性

在 开始升级 之前,请通读本指南其余部分以做好准备。

升级检查列表

警告

从 没有 v3 数据的 v2 迁移 时,如果 etcd 从现有快照恢复,但不存在 v3 ETCD_DATA_DIR/member/snap/db 文件,etcd v3.2+ 服务器会发生崩溃。这种情况出现在服务器由 v2 迁移且此前没有 v3 数据时。此限制也可防止意外丢失 v3 数据(例如 db 文件可能已被移动)。etcd 要求 v3 迁移后的操作必须有 v3 数据。v3.0 服务器包含 v3 数据之前,请勿升级到更新的 v3 版本。

监控

以下来自 v3.0.x 的指标已弃用,建议改用 go-grpc-prometheus :

  • etcd_grpc_requests_total
  • etcd_grpc_requests_failed_total
  • etcd_grpc_active_streams
  • etcd_grpc_unary_requests_duration_seconds

升级要求

要将现有 etcd 部署升级至 3.1 版本,运行中的集群版本必须为 3.0 或更高。若版本低于 3.0,请先 升级至 3.0 ,再升级至 3.1。

此外,为确保滚动升级顺利进行,运行中的集群必须处于健康状态。在继续操作前,请使用 etcdctl endpoint health 命令检查集群健康状况。

准备

在升级 etcd 之前,请务必在预发环境中测试依赖 etcd 的服务,再将升级部署到生产环境。

升级前,请对 etcd 数据执行 备份 etcd 数据 。若升级过程中出现异常,可使用此备份将系统 降级 至现有 etcd 版本。请注意,snapshot命令仅备份 v3 数据。如需备份 v2 数据,请参阅 备份 v2 数据存储 。

混合版本

升级过程中,etcd 集群支持不同版本的 etcd 成员共存,并以最低公共版本的协议运行。只有当集群中所有成员均升级至 3.1 版本后,才认为集群已完成升级。内部机制上,etcd 成员之间会相互协商以确定集群的整体版本,该版本控制报告的版本号以及所支持的功能。

限制

请注意:如果集群仅包含 v3 数据且无 v2 数据,则不受此限制影响。

如果集群正在服务的数据集大小超过 50MB,每个新升级的成员可能需要最多 2 分钟才能追上现有集群。请检查最近快照的大小以估算总数据量。换句话说,升级每个成员之间应至少等待 2 分钟。

对于数据总量更大(例如 100MB 或更多)的情况,此一次性操作可能需要更长时间。对于规模达到此类程度的大型 etcd 集群,系统管理员可在升级前自由联系 etcd 团队 ,我们将乐意提供升级流程方面的建议。

降级

如果所有成员均已升级至 v3.1 版本,集群将升级至 v3.1 版本,从该完成状态回退不可行。然而,若任一成员仍为 v3.0 版本,则集群及其操作仍保持 “v3.0” 状态,此时可从该混合集群状态恢复至所有成员均使用 v3.0 etcd 二进制文件。

请注意,务必对所有 etcd 成员的数据目录 backup the data directory 进行备份,以确保在集群完全升级后仍可执行降级操作。

升级流程

本示例演示如何升级在本地计算机上运行的 3 个成员的 v3.0 etcd 集群。

1. 检查升级要求

集群是否健康且运行 v3.0.x 版本?

$ ETCDCTL_API=3 etcdctl endpoint health --endpoints=localhost:2379,localhost:22379,localhost:32379
localhost:2379 is healthy: successfully committed proposal: took = 6.600684ms
localhost:22379 is healthy: successfully committed proposal: took = 8.540064ms
localhost:32379 is healthy: successfully committed proposal: took = 8.763432ms

$ curl http://localhost:2379/version
{"etcdserver":"3.0.16","etcdcluster":"3.0.0"}

2. 停止现有 etcd 进程

当每个 etcd 进程停止时,集群中的其他成员会记录预期的错误。这是正常的,因为集群成员之间的连接已(暂时)中断:

2017-01-17 09:34:18.352662 I | raft: raft.node: 1640829d9eea5cfb elected leader 1640829d9eea5cfb at term 5
2017-01-17 09:34:18.359630 W | etcdserver: failed to reach the peerURL(http://localhost:2380) of member fd32987dcd0511e0 (Get http://localhost:2380/version: dial tcp 127.0.0.1:2380: getsockopt: connection refused)
2017-01-17 09:34:18.359679 W | etcdserver: cannot get the version of member fd32987dcd0511e0 (Get http://localhost:2380/version: dial tcp 127.0.0.1:2380: getsockopt: connection refused)
2017-01-17 09:34:18.548116 W | rafthttp: lost the TCP streaming connection with peer fd32987dcd0511e0 (stream Message writer)
2017-01-17 09:34:19.147816 W | rafthttp: lost the TCP streaming connection with peer fd32987dcd0511e0 (stream MsgApp v2 writer)
2017-01-17 09:34:34.364907 W | etcdserver: failed to reach the peerURL(http://localhost:2380) of member fd32987dcd0511e0 (Get http://localhost:2380/version: dial tcp 127.0.0.1:2380: getsockopt: connection refused)

此时建议 备份 etcd 数据 ,以便在出现任何问题时提供回退路径:

$ etcdctl snapshot save backup.db

3. 直接替换 etcd v3.1 二进制文件并启动新 etcd 进程

新版 v3.1 etcd 将向集群发布其信息:

2017-01-17 09:36:00.996590 I | etcdserver: published {Name:my-etcd-1 ClientURLs:[http://localhost:2379]} to cluster 46bc3ce73049e678

验证每个成员以及整个集群在使用新的 v3.1 etcd 二进制文件后是否恢复正常健康状态:

$ ETCDCTL_API=3 /etcdctl endpoint health --endpoints=localhost:2379,localhost:22379,localhost:32379
localhost:22379 is healthy: successfully committed proposal: took = 5.540129ms
localhost:32379 is healthy: successfully committed proposal: took = 7.321671ms
localhost:2379 is healthy: successfully committed proposal: took = 10.629901ms

升级后的成员将在整个集群完成升级前持续记录如下警告信息。这是预期行为,待所有 etcd 集群成员升级至 v3.1 后,警告将停止出现。

2017-01-17 09:36:38.406268 W | etcdserver: the local etcd version 3.0.16 is not up-to-date
2017-01-17 09:36:38.406295 W | etcdserver: member fd32987dcd0511e0 has a higher version 3.1.0
2017-01-17 09:36:42.407695 W | etcdserver: the local etcd version 3.0.16 is not up-to-date
2017-01-17 09:36:42.407730 W | etcdserver: member fd32987dcd0511e0 has a higher version 3.1.0

4. 对所有其他成员重复步骤 2 至步骤 3

5. 完成

所有成员升级完成后,集群将成功报告升级至 3.1:

2017-01-17 09:37:03.100015 I | etcdserver: updating the cluster version from 3.0 to 3.1
2017-01-17 09:37:03.104263 N | etcdserver/membership: updated the cluster version from 3.0 to 3.1
2017-01-17 09:37:03.104374 I | etcdserver/api: enabled capabilities for version 3.1
$ ETCDCTL_API=3 /etcdctl endpoint health --endpoints=localhost:2379,localhost:22379,localhost:32379
localhost:2379 is healthy: successfully committed proposal: took = 2.312897ms
localhost:22379 is healthy: successfully committed proposal: took = 2.553476ms
localhost:32379 is healthy: successfully committed proposal: took = 2.516902ms

9 - 将 etcd 从 2.3 升级到 3.0

升级 etcd 2.3 至 3.0 的流程、检查清单与注意事项

在一般情况下,从 etcd 2.3 升级到 3.0 可以实现零停机滚动升级:

  • 逐一停止 etcd v2.3 进程,并替换为 etcd v3.0 进程
  • 在所有 v3.0 进程运行后,集群即可使用 v3.0 的新特性

在 开始升级 之前,请通读本指南其余部分以做好准备。

升级检查列表

警告

从 没有 v3 数据的 v2 迁移 时,如果 etcd 从现有快照恢复,但不存在 v3 ETCD_DATA_DIR/member/snap/db 文件,etcd v3.2+ 服务器会发生崩溃。这种情况出现在服务器由 v2 迁移且此前没有 v3 数据时。此限制也可防止意外丢失 v3 数据(例如 db 文件可能已被移动)。etcd 要求 v3 迁移后的操作必须有 v3 数据。v3.0 服务器包含 v3 数据之前,请勿升级到更新的 v3 版本。

升级要求

要将现有的 etcd 部署升级至 3.0,运行中的集群版本必须为 2.3 或更高。若版本低于 2.3,请先升级至 2.3 ,再升级至 3.0。

此外,为确保滚动升级顺利进行,运行中的集群必须处于健康状态。请在继续操作前,使用 etcdctl cluster-health 命令检查集群健康状况。

准备

在升级 etcd 之前,请务必在预发环境中测试依赖 etcd 的服务,再将升级部署到生产环境。

开始前,请先 备份 etcd 数据目录 。如果升级出现问题,可以使用此备份 降级 回现有 etcd 版本。

混合版本

升级期间,etcd 集群支持不同版本的 etcd 成员共存,并以最低公共版本的协议运行。只有当集群中所有成员均升级至 3.0 版本后,才认为集群已完成升级。内部机制上,etcd 成员之间会相互协商以确定集群的整体版本,该版本控制报告的版本及支持的功能。

限制

当集群总数据量超过 50MB 时,新升级的成员可能需要最多 2 分钟才能追上现有集群。可通过检查最近快照的大小来估算总数据量。换句话说,为确保安全,应在升级每个成员之间至少等待 2 分钟。

对于数据总量更大(例如 100MB 或更多)的情况,此一次性操作可能需要更长时间。对于规模达到此类程度的大型 etcd 集群,系统管理员可在升级前自由联系 etcd 团队 ,我们将乐意提供升级流程方面的建议。

降级

如果所有成员均已升级至 v3.0,则集群将升级至 v3.0,从该完成状态回退不可行。然而,若任一成员仍为 v2.3,则集群及其操作仍处于“v2.3”状态,此时可从该混合集群状态恢复至所有成员均使用 v2.3 etcd 二进制文件。

请备份所有 etcd 成员的数据目录 ,以便在集群完成升级后仍可执行降级操作 。

升级流程

本示例详细说明如何升级运行在本地机器上的三成员 v2.3 etcd 集群。

1. 检查升级要求。

集群是否健康且运行 v.2.3.x 版本?

$ etcdctl cluster-health
member 6e3bd23ae5f1eae0 is healthy: got healthy result from http://localhost:22379
member 924e2e83e93f2560 is healthy: got healthy result from http://localhost:32379
member 8211f1d0f64f3269 is healthy: got healthy result from http://localhost:12379
cluster is healthy

$ curl http://localhost:2379/version
{"etcdserver":"2.3.x","etcdcluster":"2.3.8"}

2. 停止现有 etcd 进程

当每个 etcd 进程停止时,集群中的其他成员会记录预期的错误。这是正常的,因为集群成员之间的连接已(暂时)中断:

2016-06-27 15:21:48.624124 E | rafthttp: failed to dial 8211f1d0f64f3269 on stream Message (dial tcp 127.0.0.1:12380: getsockopt: connection refused)
2016-06-27 15:21:48.624175 I | rafthttp: the connection with 8211f1d0f64f3269 became inactive

此时建议 备份 etcd 数据目录 ,以便在出现任何问题时能够回退。

$ etcdctl backup \
      --data-dir /var/lib/etcd \
      --backup-dir /tmp/etcd_backup

3. 插入 etcd v3.0 二进制文件并启动新 etcd 进程

新版 v3.0 etcd 将向集群发布其信息:

09:58:25.938673 I | etcdserver: published {Name:infra1 ClientURLs:[http://localhost:12379]} to cluster 524400597fb1d5f6

验证每个成员以及整个集群在使用新的 v3.0 etcd 二进制文件后是否均恢复正常状态:

$ etcdctl cluster-health
member 6e3bd23ae5f1eae0 is healthy: got healthy result from http://localhost:22379
member 924e2e83e93f2560 is healthy: got healthy result from http://localhost:32379
member 8211f1d0f64f3269 is healthy: got healthy result from http://localhost:12379
cluster is healthy

升级后的成员将在整个集群完成升级前持续记录如下警告信息。这是预期行为,当所有 etcd 集群成员均升级至 v3.0 后,警告将停止出现。

2016-06-27 15:22:05.679644 W | etcdserver: the local etcd version 2.3.7 is not up-to-date
2016-06-27 15:22:05.679660 W | etcdserver: member 8211f1d0f64f3269 has a higher version 3.0.0

4. 对所有其他成员重复步骤 2 至步骤 3

5. 完成

所有成员升级完成后,集群将成功报告升级至 3.0:

2016-06-27 15:22:19.873751 N | membership: updated the cluster version from 2.3 to 3.0
2016-06-27 15:22:19.914574 I | api: enabled capabilities for version 3.0.0
$ ETCDCTL_API=3 etcdctl endpoint health
127.0.0.1:12379 is healthy: successfully committed proposal: took = 18.440155ms
127.0.0.1:32379 is healthy: successfully committed proposal: took = 13.651368ms
127.0.0.1:22379 is healthy: successfully committed proposal: took = 18.513301ms

进一步考虑事项

  • etcdctl 环境变量已更新。如果 ETCDCTL_API=2 etcdctl cluster-health 运行正常但 ETCDCTL_API=3 etcdctl endpoints health 返回 Error: grpc: timed out when dialing,请务必使用 新的变量名 。

已知问题

  • etcd < v3.1 在使用 Go > v1.7 构建时无法正常工作。详情请参见 Issue 6951 。
  • 若 etcd 服务器日志中出现 transport: http2Client.notifyError got notified that the client transport was broken unexpected EOF. 类似错误,请确保 etcd 为预构建版本,或使用以下组合构建:(etcd v3.1+ & go v1.7+) 或 (etcd <v3.1 & go v1.6.x)。
  • 在升级过程中向 v2.3 集群添加 v3 成员不被支持,可能引发 panic。详情请参见 Issue 7249 。仅在 v3 迁移期间允许混合版本的 etcd 成员。完成升级前不得进行任何成员变更操作。