postgresql - 从 Influx 迁移到 Postgres，需要提示-6ren

postgresql - 从 Influx 迁移到 Postgres，需要提示

转载作者：行者123 更新时间：2023-11-29 13:28:00

26

4

我使用 Influx 来存储我们的时间序列数据。当它工作时很酷，然后大约一个月后，它停止工作，我不明白为什么。 (类似本期https://github.com/influxdb/influxdb/issues/1386)

也许有一天 Influx 会很棒，但现在我需要使用更稳定的东西。我在考虑Postgres。我们的数据来自许多传感器，每个传感器都有一个传感器 ID。所以我正在考虑按如下方式构建我们的数据:

(pk), sensorId(string), time(timestamp), 值(float)

Influx 是为时间序列数据构建的，因此它可能有一些内置的优化。我是否需要自己进行优化以使 Postgres 高效？更具体地说，我有以下问题:

Influx 具有“系列”的概念，并且创建新系列的成本很低。所以我为每个传感器都有一个单独的系列。我应该为每个传感器创建一个单独的 Postgres 表吗？
我应该如何设置索引以加快查询速度？一个典型的查询是:选择 sensor123 最近 3 天的所有数据。
我应该为时间列使用时间戳还是整数？
如何设置保留策略？例如。自动删除超过一周的数据。
Postgres 会横向扩展吗？我可以为数据复制和负载平衡设置 ec2 集群吗？
我可以在 Postgres 中进行缩减采样吗？我在一些文章中读到我可以使用 date_trunc。但似乎我无法将它 date_trunc 到特定的时间间隔，例如25 秒。
还有其他我遗漏的注意事项吗？

提前致谢!

更新将时间列存储为大整数比将其存储为时间戳更快。我做错了什么吗？

将其存储为时间戳:

postgres=# explain analyze select * from test where sensorid='sensor_0';

Bitmap Heap Scan on test  (cost=3180.54..42349.98 rows=75352 width=25) (actual time=10.864..19.604 rows=51840 loops=1)
   Recheck Cond: ((sensorid)::text = 'sensor_0'::text)
   Heap Blocks: exact=382
   ->  Bitmap Index Scan on sensorindex  (cost=0.00..3161.70 rows=75352 width=0) (actual time=10.794..10.794 rows=51840 loops=1)
         Index Cond: ((sensorid)::text = 'sensor_0'::text)
 Planning time: 0.118 ms
 Execution time: 22.984 ms

postgres=# explain analyze select * from test where sensorid='sensor_0' and addedtime > to_timestamp(1430939804);

 Bitmap Heap Scan on test  (cost=2258.04..43170.41 rows=50486 width=25) (actual time=22.375..27.412 rows=34833 loops=1)
   Recheck Cond: (((sensorid)::text = 'sensor_0'::text) AND (addedtime > '2015-05-06 15:16:44-04'::timestamp with time zone))
   Heap Blocks: exact=257
   ->  Bitmap Index Scan on sensorindex  (cost=0.00..2245.42 rows=50486 width=0) (actual time=22.313..22.313 rows=34833 loops=1)
         Index Cond: (((sensorid)::text = 'sensor_0'::text) AND (addedtime > '2015-05-06 15:16:44-04'::timestamp with time zone))
 Planning time: 0.362 ms
 Execution time: 29.290 ms

将其存储为大整数:

postgres=# explain analyze select * from test where sensorid='sensor_0';


 Bitmap Heap Scan on test  (cost=3620.92..42810.47 rows=85724 width=25) (actual time=12.450..19.615 rows=51840 loops=1)
   Recheck Cond: ((sensorid)::text = 'sensor_0'::text)
   Heap Blocks: exact=382
   ->  Bitmap Index Scan on sensorindex  (cost=0.00..3599.49 rows=85724 width=0) (actual time=12.359..12.359 rows=51840 loops=1)
         Index Cond: ((sensorid)::text = 'sensor_0'::text)
 Planning time: 0.130 ms
 Execution time: 22.331 ms

postgres=# explain analyze select * from test where sensorid='sensor_0' and addedtime > 1430939804472;


 Bitmap Heap Scan on test  (cost=2346.57..43260.12 rows=52489 width=25) (actual time=10.113..14.780 rows=31839 loops=1)
   Recheck Cond: (((sensorid)::text = 'sensor_0'::text) AND (addedtime > 1430939804472::bigint))
   Heap Blocks: exact=235
   ->  Bitmap Index Scan on sensorindex  (cost=0.00..2333.45 rows=52489 width=0) (actual time=10.059..10.059 rows=31839 loops=1)
         Index Cond: (((sensorid)::text = 'sensor_0'::text) AND (addedtime > 1430939804472::bigint))
 Planning time: 0.154 ms
 Execution time: 16.589 ms

最佳答案

您不应该为每个传感器创建一个表。相反，您可以在表中添加一个字段来标识它属于哪个系列。您还可以有另一个表来描述有关该系列的其他属性。如果数据点可以属于多个系列，那么您将需要一个完全不同的结构。

对于您在 q2 中描述的查询，您的 recorded_at 列上的索引应该有效(time 是 sql 保留关键字，因此最好避免将其作为名称)

您应该使用 TIMESTAMP WITH TIME ZONE 作为您的时间数据类型。

保留由您决定。

Postgres 有多种分片/复制选项。这是一个很大的话题。

不确定我是否理解您对 #6 的目标，但我相信您能想出办法。

关于postgresql - 从 Influx 迁移到 Postgres，需要提示，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/30020699/

26

4

0

文章推荐： php - 如何进行查询以获取特定的帖子 ID？

文章推荐： python - Django模型查询

postgres-xl - 在单机配置上设置 Postgres-XL
我想知道这里是否有人有安装 Postgres-XL 的经验，新的开源多线程版本的 PostgreSQL。我计划将一组 1-2 TB 的数据库从常规 Postgres 9.3 迁移到 XL，并且想知道这
postgresql - postgres 备份脚本不是 postgres 用户
我想创建一个 postgres 备份脚本，但我不想使用 postgres 用户，因为我所在的 unix 系统几乎没有限制。我想要做的是在 crontab 上以 unix 系统(网络)的普通用户身份运行
node-postgres - 在 node-postgres 中使用字符串文字进行参数化查询
我正在尝试编写一个 node-postgres 查询，它采用一个整数作为参数在间隔中使用: const query = { text: `SELECT foo
postgresql - 如何从命令行停止 (postgres.app) postgres 集群
如何在不使用 gui 的情况下停止特定的 Postgres.app 集群。我想使用 bash/Terminal.app 而不是 gui 我还应该指出，Postgres 应用程序有一个这样的菜单如果
security - 在 POSTGRES 数据库中为 'postgres' 用户添加密码是个好主意吗？
关闭。这个问题是opinion-based .它目前不接受答案。想要改进这个问题？更新问题，以便 editing this post 可以用事实和引用来回答它. 关闭 9 年前。 Improve
postgresql - Postgres docker "User "postgres“没有分配密码。”
我正在使用 docker 运行 Postgres 图像。它曾经在 Windows10 和 Ubuntu 18.04 上运行没有任何问题。在 Ubuntu 系统上重新克隆项目后，它在运行 docker
python - 从另一个 postgres 表更新一个 postgres 表
我正在使用 python(比如表 A)将批处理 csv 文件加载到 postgres。我正在使用 pandas 将数据上传到更快的 block 中。 for chunk in pd.read_csv(
postgresql - 在源 Postgres 服务器离线时迁移 Postgres DB
所以是的，标题说明了一切，我需要以某种方式将 DB 从源服务器获取到新服务器，但更重要的是旧服务器正在崩溃 :P 有什么方法可以将它全部移动到新服务器并导入它？旧服务器只是拒绝再运行 Postgre
postgresql - Postgres systemd 单元文件如何确定要运行的 Postgres 版本？
这主要是出于好奇而提出的问题。我正在浏览 Postgres systemd 单元文件，以了解 systemd 可以做什么。 Postgres 有两个 systemd 单元文件。一个用于代替 syste
postgresql - 为什么 Postgres 不要求用户 postgres 的密码？
从我在 pg_hba.conf 中读到的内容，我推断，为了确保提示我输入 postgres 用户的密码，我应该从当前的“对等”编辑 pg_hba.conf 的前两个条目的方法'到'密码'或'md5'，
sql - Postgres : user mapping not found for "postgres"
我已连接到架构 apm。尝试执行函数并出现以下错误: ERROR: user mapping not found for "postgres" 数据库连接信息说: apm on postgres@
postgresql - 无法创建用户 postgres : role "postgres" does not exists
我在 ubuntu 12.04 服务器上，我正在尝试安装 postgresql。截至目前，我已成功安装它但无法配置它。我需要创建一个角色才能继续前进，我在终端中运行了这个命令: root@hostna
database - 无法以 'postgres' 用户身份登录到 'postgres' 数据库
我无法以“postgres”用户身份登录到“postgres”数据库。操作系统:REHL 服务器版本 6.3PostgreSQL 版本:8.4有一个数据库“jiradb”用作 JIRA 6.0.8 的
postgresql - 在 docker postgres 容器中导入 postgres 数据库
我正在尝试将现有数据库导入 postgres docker 容器。这就是我的处理方式: docker run --name pg-docker -e POSTGRES_PASSWORD=*****
postgresql - 重新启动 postgres 后使用 Grails 自动重新连接到 postgres
我们的 Web 应用程序在 postgres 9.3 和 Grails 2.5.3 上运行。当我们重新启动 postgres (/etc/init.d/postgresql restart) 并访问网
postgresql - 如何将现有的 postgres 数据文件夹复制和使用到 docker postgres 容器中
我想构建 postgres docker 容器来测试一些问题。我有: postgres 文件的归档文件夹(/var/lib/postgres/data/) 将文件夹放入 docker postgres
postgresql - 如何将 postgres json 列表转换为全小写的 postgres 数组
我有一个名为“stuff”的表，其中有一个名为“tags”的 json 列，用于存储标签列表，还有一个名为“id”的列，它是表中每一行的主键。我正在使用 postgres 数据库。例如，一行看起来像这
python - postgres 中的锁定机制/postgres 中的死锁。 [我正在使用 sqlalchemy]
我对 sqlalchemy-psql 中的锁定机制是如何工作的感到非常困惑。我正在运行一个带有 sqlalchemy 和 postgres 的 python-flask 应用程序。由于我有多个线程处理
sql - Postgres 函数比查询/Postgres 8.4 慢
我(必须)使用 Postgres 8.4 数据库。在这个数据库中，我创建了一个函数: CREATE OR REPLACE FUNCTION counter (mindate timestamptz,m
linux - Postgres : psql: FATAL: role "postgres" does not exist
我已经使用 PostgreSQL 几天了，它运行良好。我一直在通过默认的 postgres 数据库用户和另一个具有权限的用户使用它。今天中午(在一切正常之后)它停止工作，我再也无法回到数据库中。我会

首页

博学

6Ren·AI

商城

postgresql - 从 Influx 迁移到 Postgres，需要提示