postgresql - pgpool-ii 主/从模式 : How can I most easily trigger a failover?-6ren

postgresql - pgpool-ii 主/从模式 : How can I most easily trigger a failover?

转载作者：行者123 更新时间：2023-11-29 13:53:55

因此，我正在使用一些本地虚拟机测试一些玩具 postgresql 基础设施，以确定 pgpool 在故障转移时的行为方式。我已经配置了一个基本设置，其中我有两台数据库机器(192.168.0.2 和 192.168.0.3)和一台 pgpool 机器(192.168.0.4)。 192.168.0.3 已使用流复制设置为 192.168.0.2 的从站。 pgpool-ii 已使用以下配置:

listen_addresses = '*'
backend_hostname0 = '192.168.0.2'
backend_port0 = 5432
backend_weight0 = 1
backend_data_directory0 = '/var/lib/postgresql/9.4/main/'
backend_flag0 = 'ALLOW_TO_FAILOVER'
backend_hostname1 = '192.168.0.3'
backend_port1 = 5432
backend_weight1 = 1
backend_data_directory1 = '/var/lib/postgresql/9.4/main/'
backend_flag1 = 'ALLOW_TO_FAILOVER'
enable_pool_hba = on
replication_mode = false
master_slave_mode = on
master_slave_sub_mode = 'stream'
fail_over_on_backend_error = true
failover_command = '/root/pgpool_failover_stream.sh %d %H /tmp/postgresql.trigger.5432'
load_balance_mode = false

我已经确认一切正常。也就是说，当我更改主数据库时，复制正在运行，我可以使用示例应用程序连接到主数据库、从数据库和 pgpool-ii，并获得我期望的结果。

现在，我启动了一个连接到 pgpool 的长时间运行的应用程序，然后尝试通过 SSH 连接到主数据库服务器并强制结束 postgres 任务(service postgresql stop 以 root 身份)来引起故障转移。我的应用程序继续正确执行查询，但没有发生故障转移(脚本尚未运行)。我什至测试过直接连接到 master 数据库，当我停止 postgres 服务时，我最终导致应用程序崩溃。

我做错了什么吗？我没有正确配置我的 pgpool 吗？还是有更好的方法来触发故障转移？

编辑:根据要求，这是发生第一个错误的日志部分:

...
2016-03-15 18:47:15: pid 1232: DEBUG:  initializing backend status
2016-03-15 18:47:15: pid 1231: DEBUG:  initializing backend status
2016-03-15 18:47:15: pid 1230: DEBUG:  initializing backend status
2016-03-15 18:47:15: pid 1209: ERROR:  failed to authenticate
2016-03-15 18:47:15: pid 1209: DETAIL:  invalid authentication message response type, Expecting 'R' and received 'E'
2016-03-15 18:47:15: pid 1209: LOG:  find_primary_node: checking backend no 1
2016-03-15 18:47:15: pid 1209: ERROR:  failed to authenticate
2016-03-15 18:47:15: pid 1209: DETAIL:  invalid authentication message response type, Expecting 'R' and received 'E'
2016-03-15 18:47:15: pid 1209: DEBUG:  find_primary_node: no primary node found
...

奇怪的是，我仍然可以连接到 pgpool 并执行查询，所以显然我不明白那里的东西。

编辑 2:这些是我在 master 上 service postgresql shutdown 后得到的错误。我显示所有内容以开始关闭 pgpool。

...
2016-03-16 17:24:57: pid 1012: DEBUG:  session context: clearing doing extended query messaging. DONE
2016-03-16 17:24:57: pid 1012: DEBUG:  session context: setting doing extended query messaging. DONE
2016-03-16 17:24:57: pid 1012: DEBUG:  session context: setting query in progress. DONE
2016-03-16 17:24:57: pid 1012: DEBUG:  reading backend data packet kind
2016-03-16 17:24:57: pid 1012: DETAIL:  backend:0 of 2 kind = 'E'
2016-03-16 17:24:57: pid 1012: DEBUG:  processing backend response
2016-03-16 17:24:57: pid 1012: DETAIL:  received kind 'E'(45) from backend
2016-03-16 17:24:57: pid 1012: ERROR:  unable to forward message to frontend
2016-03-16 17:24:57: pid 1012: DETAIL:  FATAL error occured on backend
2016-03-16 17:24:57: pid 1012: DEBUG:  session context: setting query in progress. DONE
2016-03-16 17:24:57: pid 1012: DEBUG:  decide where to send the queries
2016-03-16 17:24:57: pid 1012: DETAIL:  destination = 3 for query= "DISCARD ALL"
2016-03-16 17:24:57: pid 1012: DEBUG:  waiting for query response
2016-03-16 17:24:57: pid 1012: DETAIL:  waiting for backend:0 to complete the query
2016-03-16 17:24:57: pid 1012: FATAL:  unable to read data from DB node 0
2016-03-16 17:24:57: pid 1012: DETAIL:  EOF encountered with backend
2016-03-16 17:24:57: pid 998: DEBUG:  reaper handler
2016-03-16 17:24:57: pid 998: LOG:  child process with pid: 1012 exits with status 256
2016-03-16 17:24:57: pid 998: LOG:  fork a new child process with pid: 1033
2016-03-16 17:24:57: pid 998: DEBUG:  reaper handler: exiting normally
2016-03-16 17:24:57: pid 1033: DEBUG:  initializing backend status
2016-03-16 17:25:02: pid 1031: DEBUG:  PCP child receives shutdown request signal 2
2016-03-16 17:25:02: pid 1029: LOG:  child process received shutdown request signal 2
...

请注意，当主服务器关闭时，我的示例应用程序确实死掉了。

编辑 3:在正确设置 sr_check_period、sr_check_user、sr_check_password 后，我在新日志中遇到错误，所有以前的错误现在都消失了:

2016-03-31 17:45:00: pid 18363: DEBUG:  detect error: kind: 1
2016-03-31 17:45:00: pid 18363: DEBUG:  reading backend data packet kind
2016-03-31 17:45:00: pid 18363: DETAIL:  backend:0 of 2 kind = '1'
...
2016-03-31 17:45:00: pid 18363: DEBUG:  detect error: kind: S

最佳答案

故障转移脚本未执行的原因可能有多种。主要步骤是启用系统日志的 log_destination 属性并启用 Debug模式 (debug_level =1)。

我亲眼目睹了故障转移脚本无法获取 %d、%H(特殊字符)参数的情况，因此脚本无法通过 ssh 连接到从属服务器并访问触发器文件。

如果您发布相同的日志文件，我可以提供更多详细信息。

基于新日志:我可以看到 ERROR: failed to authenticate 。能否检查下pgpool的以下参数是否配置正确

健康检查用户
健康检查密码
恢复用户
恢复密码
wd_lifecheck_user
wd_lifecheck_password
sr_check_user
sr_check_password

我希望你已经按照更改postgres用户密码的步骤进行操作

alter user postgres password 'yourpassword'

并确保您在所有情况下都提供相同的密码。

从日志来看，这看起来像是身份验证问题。你能告诉我你使用的 pgpool 版本吗？

这些是我们用于设置 3 台机器(1 台主机、1 台从机和 1 台机器用于 pgpool)的配置我已经修改为适合你的ip地址

 listen_addresses = '*'
  port = 5433
  socket_dir = '/var/run/postgresql'
  pcp_port = 9898
  pcp_socket_dir = '/var/run/postgresql'

  backend_hostname0 = '192.168.0.2'
  backend_port0 = 5432
  backend_weight0 = 1
  backend_data_directory0 = '/var/lib/postgresql/9.4/main'
  backend_flag0 = 'ALLOW_TO_FAILOVER'

  backend_hostname1 = '192.168.0.3'
  backend_port1 = 5432
  backend_weight1 = 1
  backend_data_directory1 = '/var/lib/postgresql/9.4/main'
  backend_flag1 = 'ALLOW_TO_FAILOVER'

  enable_pool_hba = on
  pool_passwd = ''
  authentication_timeout = 60
  ssl = off
  num_init_children = 4
  max_pool = 2
  child_life_time = 300 
  child_max_connections = 0
  connection_life_time = 0
  client_idle_limit = 0
  log_destination = 'stderr,syslog'
  print_timestamp = on
  log_connections = on
  log_hostname = on
  log_statement = on
  log_per_node_statement = on
  log_standby_delay = 'none'
  syslog_facility = 'LOCAL0'
  syslog_ident = 'pgpool'
  debug_level = 1
  pid_file_name = '/var/run/postgresql/pgpool.pid'
  logdir = '/var/log/postgresql'
  connection_cache = on
  reset_query_list = 'ABORT; DISCARD ALL'

  replication_mode = off
  replicate_select = off
  insert_lock = on
  lobj_lock_table = ''
  replication_stop_on_mismatch = off
  failover_if_affected_tuples_mismatch = off

  load_balance_mode = off
  ignore_leading_white_space = on
  white_function_list = ''
  black_function_list = 'nextval,setval'

  master_slave_mode = on
  master_slave_sub_mode = 'stream'
  sr_check_period = 10
  sr_check_user = 'postgres'
  sr_check_password = 'postgres123'
  delay_threshold = 0
  follow_master_command = ''
  parallel_mode = off
  pgpool2_hostname = 'pgmaster'

  system_db_hostname  = 'localhost'
  system_db_port = 5432
  system_db_dbname = 'pgpool'
  system_db_schema = 'pgpool_catalog'
  system_db_user = 'pgpool'
  system_db_password = ''

  health_check_period = 5
  health_check_timeout = 20
  health_check_user = 'postgres'
  health_check_password = 'postgres123'
  health_check_max_retries = 2
  health_check_retry_delay = 1

  failover_command = '/usr/sbin/failover_modified.sh %d "%H" %P /var/lib/postgresql/9.4/main/pgsql.trigger.5432'
  failback_command = ''
  fail_over_on_backend_error = on
  search_primary_node_timeout = 10

  recovery_user = 'postgres'
  recovery_password = 'postgres123'
  recovery_1st_stage_command = ''
  recovery_2nd_stage_command = ''
  recovery_timeout = 90
  client_idle_limit_in_recovery = 0

  use_watchdog = off
  trusted_servers = ''
  ping_path = '/bin'
  wd_hostname = ''
  wd_port = 9000
  wd_authkey = ''
  delegate_IP = ''
  ifconfig_path = '/sbin'
  if_up_cmd = 'ifconfig eth0:0 inet $_IP_$ netmask 255.255.255.0'
  if_down_cmd = 'ifconfig eth0:0 down'
  arping_path = '/usr/sbin'  
  arping_cmd = 'arping -U $_IP_$ -w 1'

  clear_memqcache_on_escalation = on
  wd_escalation_command = ''

  wd_lifecheck_method = 'heartbeat'
  wd_interval = 10
  wd_heartbeat_port = 9694
  wd_heartbeat_keepalive = 2
  wd_heartbeat_deadtime = 30
  heartbeat_destination0 = '192.168.0.2'
  heartbeat_destination_port0 = 9694
  heartbeat_device0 = ''

  heartbeat_destination1 = '192.168.0.3'
  wd_life_point = 3
  wd_lifecheck_query = 'SELECT 1'
  wd_lifecheck_dbname = 'postgres'
  wd_lifecheck_user = 'postgres'
  wd_lifecheck_password = 'postgres123'

  relcache_expire = 0
  relcache_size = 256
  check_temp_table = on

  memory_cache_enabled = off
  memqcache_method = 'shmem'
  memqcache_memcached_host = 'localhost'
  memqcache_memcached_port = 11211
  memqcache_total_size = 67108864
  memqcache_max_num_cache = 1000000
  memqcache_expire = 0
  memqcache_auto_cache_invalidation = on
  memqcache_maxcache = 409600
  memqcache_cache_block_size = 1048576
  memqcache_oiddir = '/var/log/pgpool/oiddir'
  white_memqcache_table_list = ''
  black_memqcache_table_list = ''

此外，我希望您已经修改了 pool_hba.conf 以启用对 master 和 slave 的访问

关于postgresql - pgpool-ii 主/从模式 : How can I most easily trigger a failover?，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/35735353/

文章推荐： postgresql - Postgres 表似乎部分丢失

文章推荐： php - 多个 MySQL 查询还是一个具有多个 JOIN 的查询？

文章推荐：使用名称而不是问号的 SQL 参数化查询

database - 主-主 vs 主-从数据库架构？
我听说过两种数据库架构。大师级主从 master-master不是更适合现在的web吗，因为它就像Git一样，每个单元都有整套数据，如果一个宕机也无所谓。主从让我想起了 SVN(我不喜欢它)，你
mysql 复制主-主-主
我们当前将 MySQL 配置为支持故障转移:Site1 Site2。当它们被设置为主/主时。在给定时间点，应用程序服务器仅主动写入一个站点。我们想要设置一个新的故障转移站点。然后我们将拥有 Site
database - 主-主 vs 主从数据库架构？
我听说过两种数据库架构。大师-大师主从 master-master 不是更适合当今的网络吗，因为它就像 Git，每个单元都有整套数据，如果其中一个发生故障，也没关系。主从让我想起 SVN(我不喜
MySQL:类别和子类别的交叉表引用(主、主、名称)？
我正在创建一个标记为类别的表，其中主类别(父列)包含 0，子类别包含父类别的 ID。我听说这叫引用。我的问题:这张表的结构正确吗？或者是否有更好的方法，例如实现遍历树或类似方法？ CREATE TAB
C++ 主 vs C 主
我正在阅读一份关于 C++ 与 C 的文档。该文档说与 C 相比，C++ 编写得非常紧凑。一个例子是，C 允许 main() 函数类型为 void。另一方面，C++ 不允许这样做，他给出了标准中的以下
java - C 主 vs Java 主
C main函数和Java main函数有什么区别？ int main( int argc, const char* argv[] ) 对比 public static void main(Strin
responsive-design - 3 列(左、主、右)站点，但较小的设备首先是主(主、左、右)
我一直摸不着头脑，但运气不好。设计器有一个包含 3 栏的站点、两个侧边栏和一个主要内容区域。专为桌面设计，左栏、主要内容、右栏。但是，在较小的设备上，我们希望首先堆叠主要内容。所以通常情况下，你可
Jenkins 主/从配置
我一直在阅读有关 Jenkins 主/从配置的信息，但我仍然有一些问题: 是不是真的没有像 Jenkins 主站那样安装和启动从站 Jenkins？我假设我会以相同的方式安装一个主 Jenkins 和
c# - 主/细节的通用Viewmodel模式
据我了解，Viemodel中MVVM背后的概念包括业务逻辑和/或诸如暴露于 View 的数据的主/明细关系之类的事物因此，正如我发现的那样，有很多ORM生成器，例如模型的telerik a.o以及另
elasticsearch - 主/副本计分不一致
我们有一个群集，其中包含3个主分区，每个主分区有2个副本。主/副本分片的总文档数相同；但是，对于同一查询/文档，我们得到3个不同的分数。当我们将preference = primary添加为查询参数时
java - 从指定目录启动Java应用程序(主)
我有一个非常大/旧/长时间运行的项目，它使用相对于启动目录的路径访问文件资源(即应用程序仅在从特定目录启动时才工作)。当我需要调试程序时，我可以从 eclipse 启动它并使用“运行配置”->->“工
c - 主/从设备出现段错误
谁能向我解释一下为什么我在这段代码上遇到段错误？我一直试图弄清楚这一点，但在各种搜索中却一无所获。当我运行代码而不调用 main(argc, argv) 时，它会运行。 Slave 仅将 argv 中
ios - 主 - 详细项目未按预期工作
使用 xcode 中的默认项目作为主从应用程序，如果我在折叠委托(delegate)中放置 print 调试语句，当我旋转设备时它似乎永远不会被触发(事实上我永远无法触发它)。我编辑的代码位于 Ap
mysql 主/从故障转移
是否有任何产品可以使 mysql 主/从故障转移过程更容易？一些可以自动发生的事情，而不是手动修复它。最佳答案 [...稍后...;) 你所说的“更容易”是什么？MySQL 有很多解决方案: MyS
MySQL 主/主复制仅以一种方式工作
我有两个 mysql 数据库。我想做主/主复制。复制以一种方式进行。然而，反过来说却不然。该错误表明它无法与用户“test@IPADDRESS”连接。如何将用户名更改为 repl？从未进行过测试，
MySQL 主-主复制
我正在尝试在 MySQL 中运行以下查询: GRANT REPLICATION SLAVE ON *.* TO 'replication'@’10.141.2.%’ IDENTIFIED BY ‘sl
android - 主/从流程中的操作栏菜单项
我正在尝试使用 Android 提供的主/详细流程模板创建一个应用程序，并且我正在尝试将多个操作栏菜单项添加到操作栏的主要部分和详细信息部分。这就是我要实现的目标: (来源:softwarecrew.
C++ 主/工
我正在寻找一个跨平台的 C++ master/worker 库或工作队列库。一般的想法是我的应用程序将创建某种任务或工作对象，将它们传递给工作主机或工作队列，这将依次在单独的线程或进程中执行工作。为了
MySQL 主/外键大小？
我似乎看到很多人在他们的 MySQL 模式中任意分配大尺寸的主/外键字段，例如 INT(11) 甚至 WordPress 使用的 BIGINT(20)。如果我错了，请纠正我，但即使是 INT(4)
sql - 主/外键与没有主键的单表的性能
如果我有一个可以与多个键相关联的用户，正确的表设置应该是: 一个表有两列，例如: UserName | Key 没有主键且用户可以有多行，或者: 具有匹配标识符的两个表 Table 1 Us

行者123

个人简介

我是一名优秀的程序员,十分优秀！

作者热门文章

滴滴打车优惠券免费领取

全站热门文章

首页

博学

6Ren·AI

商城

postgresql - pgpool-ii 主/从模式 : How can I most easily trigger a failover?