Upgrade etcd from 3.3 to 3.4
In the general case, upgrading from etcd 3.3 to 3.4 can be a zero-downtime, rolling upgrade:
- one by one, stop the etcd v3.3 processes and replace them with etcd v3.4 processes
- after running all v3.4 processes, new features in v3.4 are available to the cluster
Before starting an upgrade , read through the rest of this guide to prepare.
Upgrade checklists
When migrating from v2 with no v3 data
, etcd server v3.2+ panics when etcd restores from existing snapshots but no v3 ETCD_DATA_DIR/member/snap/db file. This happens when the server had migrated from v2 with no previous v3 data. This also prevents accidental v3 data loss (e.g. db file might have been moved). etcd requires that post v3 migration can only happen with v3 data. Do not upgrade to newer v3 versions until v3.0 server contains v3 data.
Highlighted breaking changes in 3.4.
Make ETCDCTL_API=3 etcdctl default
ETCDCTL_API=3 is now the default.
Make etcd --enable-v2=false default
etcd --enable-v2=false
is now the default.
This means, unless etcd --enable-v2=true is specified, etcd v3.4 server would not serve v2 API requests.
If v2 API were used, make sure to enable v2 API in v3.4:
Other HTTP APIs will still work (e.g. [CLIENT-URL]/metrics, [CLIENT-URL]/health, v3 gRPC gateway).
Deprecated etcd --ca-file and etcd --peer-ca-file flags
--ca-file and --peer-ca-file flags are deprecated; they have been deprecated since v2.1.
Note setting this parameter will also automatically enable client cert authentication no matter what value is set for --client-cert-auth.
Deprecated grpc.ErrClientConnClosing error
grpc.ErrClientConnClosing has been deprecated in gRPC >= 1.10
.
Require grpc.WithBlock for client dial
The new client balancer
uses an asynchronous resolver to pass endpoints to the gRPC dial function. As a result, v3.4 client requires grpc.WithBlock dial option to wait until the underlying connection is up.
Deprecating etcd_debugging_mvcc_db_total_size_in_bytes Prometheus metrics
v3.4 promotes etcd_debugging_mvcc_db_total_size_in_bytes Prometheus metrics to etcd_mvcc_db_total_size_in_bytes, in order to encourage etcd storage monitoring.
etcd_debugging_mvcc_db_total_size_in_bytes is still served in v3.4 for backward compatibilities. It will be completely deprecated in v3.5.
Note that etcd_debugging_* namespace metrics have been marked as experimental. As we improve monitoring guide, we may promote more metrics.
Deprecating etcd_debugging_mvcc_put_total Prometheus metrics
v3.4 promotes etcd_debugging_mvcc_put_total Prometheus metrics to etcd_mvcc_put_total, in order to encourage etcd storage monitoring.
etcd_debugging_mvcc_put_total is still served in v3.4 for backward compatibilities. It will be completely deprecated in v3.5.
Note that etcd_debugging_* namespace metrics have been marked as experimental. As we improve monitoring guide, we may promote more metrics.
Deprecating etcd_debugging_mvcc_delete_total Prometheus metrics
v3.4 promotes etcd_debugging_mvcc_delete_total Prometheus metrics to etcd_mvcc_delete_total, in order to encourage etcd storage monitoring.
etcd_debugging_mvcc_delete_total is still served in v3.4 for backward compatibilities. It will be completely deprecated in v3.5.
Note that etcd_debugging_* namespace metrics have been marked as experimental. As we improve monitoring guide, we may promote more metrics.
Deprecating etcd_debugging_mvcc_txn_total Prometheus metrics
v3.4 promotes etcd_debugging_mvcc_txn_total Prometheus metrics to etcd_mvcc_txn_total, in order to encourage etcd storage monitoring.
etcd_debugging_mvcc_txn_total is still served in v3.4 for backward compatibilities. It will be completely deprecated in v3.5.
Note that etcd_debugging_* namespace metrics have been marked as experimental. As we improve monitoring guide, we may promote more metrics.
Deprecating etcd_debugging_mvcc_range_total Prometheus metrics
v3.4 promotes etcd_debugging_mvcc_range_total Prometheus metrics to etcd_mvcc_range_total, in order to encourage etcd storage monitoring.
etcd_debugging_mvcc_range_total is still served in v3.4 for backward compatibilities. It will be completely deprecated in v3.5.
Note that etcd_debugging_* namespace metrics have been marked as experimental. As we improve monitoring guide, we may promote more metrics.
Deprecating etcd --log-output flag (now --log-outputs)
Rename etcd --log-output to --log-outputs
to support multiple log outputs. etcd --logger=capnslog does not support multiple log outputs.
etcd --log-output will be deprecated in v3.5. etcd --logger=capnslog will be deprecated in v3.5.
v3.4 adds etcd --logger=zap --log-outputs=stderr support for structured logging and multiple log outputs. Main motivation is to promote automated etcd monitoring, rather than looking back server logs when it starts breaking. Future development will make etcd log as few as possible, and make etcd easier to monitor with metrics and alerts. etcd --logger=capnslog will be deprecated in v3.5.
Changed log-outputs field type in etcd --config-file to []string
Now that log-outputs (old field name log-output) accepts multiple writers, etcd configuration YAML file log-outputs field must be changed to []string type as below:
Renamed embed.Config.LogOutput to embed.Config.LogOutputs
Renamed embed.Config.LogOutput to embed.Config.LogOutputs
to support multiple log outputs. And changed embed.Config.LogOutput type from string to []string
to support multiple log outputs.
v3.5 deprecates capnslog
v3.5 will deprecate etcd --log-package-levels flag for capnslog; etcd --logger=zap --log-outputs=stderr will the default. v3.5 will deprecate [CLIENT-URL]/config/local/log endpoint.
Deprecating etcd --debug flag (now --log-level=debug)
v3.4 deprecates etcd --debug
flag. Instead, use etcd --log-level=debug flag.
Deprecated pkg/transport.TLSInfo.CAFile field
Deprecated pkg/transport.TLSInfo.CAFile field.
Changed embed.Config.SnapCount to embed.Config.SnapshotCount
To be consistent with the flag name etcd --snapshot-count, embed.Config.SnapCount field has been renamed to embed.Config.SnapshotCount:
Changed etcdserver.ServerConfig.SnapCount to etcdserver.ServerConfig.SnapshotCount
To be consistent with the flag name etcd --snapshot-count, etcdserver.ServerConfig.SnapCount field has been renamed to etcdserver.ServerConfig.SnapshotCount:
Changed function signature in package wal
Changed wal function signatures to support structured logger.
Changed IntervalTree type in package pkg/adt
pkg/adt.IntervalTree is now defined as an interface.
Deprecated embed.Config.SetupLogging
embed.Config.SetupLogging has been removed in order to prevent wrong logging configuration, and now set up automatically.
Changed gRPC gateway HTTP endpoints (replaced /v3beta with /v3)
Before
After
Requests to /v3beta endpoints will redirect to /v3, and /v3beta will be removed in 3.5 release.
Deprecated container image tags
latest and minor version images tags are deprecated:
Server upgrade checklists
Upgrade requirements
To upgrade an existing etcd deployment to 3.4, the running cluster must be 3.3 or greater. If it’s before 3.3, please upgrade to 3.3 before upgrading to 3.4.
Also, to ensure a smooth rolling upgrade, the running cluster must be healthy. Check the health of the cluster by using the etcdctl endpoint health command before proceeding.
Preparation
Before upgrading etcd, always test the services relying on etcd in a staging environment before deploying the upgrade to the production environment.
Before beginning, download the snapshot backup
. Should something go wrong with the upgrade, it is possible to use this backup to downgrade
back to existing etcd version. Please note that the snapshot command only backs up the v3 data. For v2 data, see backing up v2 datastore
.
Mixed versions
While upgrading, an etcd cluster supports mixed versions of etcd members, and operates with the protocol of the lowest common version. The cluster is only considered upgraded once all of its members are upgraded to version 3.4. Internally, etcd members negotiate with each other to determine the overall cluster version, which controls the reported version and the supported features.
Limitations
Note: If the cluster only has v3 data and no v2 data, it is not subject to this limitation.
If the cluster is serving a v2 data set larger than 50MB, each newly upgraded member may take up to two minutes to catch up with the existing cluster. Check the size of a recent snapshot to estimate the total data size. In other words, it is safest to wait for 2 minutes between upgrading each member.
For a much larger total data size, 100MB or more , this one-time process might take even more time. Administrators of very large etcd clusters of this magnitude can feel free to contact the etcd team before upgrading, and we’ll be happy to provide advice on the procedure.
Downgrade
If all members have been upgraded to v3.4, the cluster will be upgraded to v3.4, and downgrade from this completed state is not possible. If any single member is still v3.3, however, the cluster and its operations remains “v3.3”, and it is possible from this mixed cluster state to return to using a v3.3 etcd binary on all members.
Please download the snapshot backup to make downgrading the cluster possible even after it has been completely upgraded.
Upgrade procedure
This example shows how to upgrade a 3-member v3.3 etcd cluster running on a local machine.
Step 1: check upgrade requirements
Is the cluster healthy and running v3.3.x?
Step 2: download snapshot backup from leader
Download the snapshot backup to provide a downgrade path should any problems occur.
etcd leader is guaranteed to have the latest application data, thus fetch snapshot from leader:
Step 3: stop one existing etcd server
When each etcd process is stopped, expected errors will be logged by other cluster members. This is normal since a cluster member connection has been (temporarily) broken:
Step 4: restart the etcd server with same configuration
Restart the etcd server with same configuration but with the new etcd binary.
The new v3.4 etcd will publish its information to the cluster. At this point, cluster still operates as v3.3 protocol, which is the lowest common version.
{"level":"info","ts":1526586617.1647713,"caller":"membership/cluster.go:485","msg":"set initial cluster version","cluster-id":"7dee9ba76d59ed53","local-member-id":"7339c4e5e833c029","cluster-version":"3.0"}
{"level":"info","ts":1526586617.1648536,"caller":"api/capability.go:76","msg":"enabled capabilities for version","cluster-version":"3.0"}
{"level":"info","ts":1526586617.1649303,"caller":"membership/cluster.go:473","msg":"updated cluster version","cluster-id":"7dee9ba76d59ed53","local-member-id":"7339c4e5e833c029","from":"3.0","from":"3.3"}
{"level":"info","ts":1526586617.1649797,"caller":"api/capability.go:76","msg":"enabled capabilities for version","cluster-version":"3.3"}
{"level":"info","ts":1526586617.2107732,"caller":"etcdserver/server.go:1770","msg":"published local member to cluster through raft","local-member-id":"7339c4e5e833c029","local-member-attributes":"{Name:s1 ClientURLs:[http://localhost:2379]}","request-path":"/0/members/7339c4e5e833c029/attributes","cluster-id":"7dee9ba76d59ed53","publish-timeout":7}
Verify that each member, and then the entire cluster, becomes healthy with the new v3.4 etcd binary:
Un-upgraded members will log warnings like the following until the entire cluster is upgraded.
This is expected and will cease after all etcd cluster members are upgraded to v3.4:
Step 5: repeat step 3 and step 4 for rest of the members
When all members are upgraded, the cluster will report upgrading to 3.4 successfully:
Member 1:
{"level":"info","ts":1526586949.0920913,"caller":"api/capability.go:76","msg":"enabled capabilities for version","cluster-version":"3.4"}{"level":"info","ts":1526586949.0921566,"caller":"etcdserver/server.go:2272","msg":"cluster version is updated","cluster-version":"3.4"}
Member 2:
{"level":"info","ts":1526586949.092117,"caller":"membership/cluster.go:473","msg":"updated cluster version","cluster-id":"7dee9ba76d59ed53","local-member-id":"729934363faa4a24","from":"3.3","from":"3.4"}{"level":"info","ts":1526586949.0923078,"caller":"api/capability.go:76","msg":"enabled capabilities for version","cluster-version":"3.4"}
Member 3:
{"level":"info","ts":1526586949.0921423,"caller":"membership/cluster.go:473","msg":"updated cluster version","cluster-id":"7dee9ba76d59ed53","local-member-id":"b548c2511513015","from":"3.3","from":"3.4"}{"level":"info","ts":1526586949.0922918,"caller":"api/capability.go:76","msg":"enabled capabilities for version","cluster-version":"3.4"}