c++ - __rdtscp calibration unstable under Linux on Intel Xeon X5550-6ren

c++ - __rdtscp calibration unstable under Linux on Intel Xeon X5550

转载作者：太空狗更新时间：2023-10-29 11:36:12

24

4

我正在尝试使用 __rdtscp 内部函数来测量时间间隔。目标平台是 Linux x64，CPU Intel Xeon X5550。尽管为该处理器设置了 constant_tsc 标志，但校准 __rdtscp 会给出截然不同的结果:

$ taskset -c 1 ./ticks
Ticks per usec: 256
$ taskset -c 1 ./ticks
Ticks per usec: 330.667
$ taskset -c 1 ./ticks
Ticks per usec: 345.043
$ taskset -c 1 ./ticks
Ticks per usec: 166.054
$ taskset -c 1 ./ticks
Ticks per usec: 256
$ taskset -c 1 ./ticks
Ticks per usec: 345.043
$ taskset -c 1 ./ticks
Ticks per usec: 256
$ taskset -c 1 ./ticks
Ticks per usec: 330.667
$ taskset -c 1 ./ticks
Ticks per usec: 256
$ taskset -c 1 ./ticks
Ticks per usec: 330.667
$ taskset -c 1 ./ticks
Ticks per usec: 330.667
$ taskset -c 1 ./ticks
Ticks per usec: 345.043
$ taskset -c 1 ./ticks
Ticks per usec: 256
$ taskset -c 1 ./ticks
Ticks per usec: 125.388
$ taskset -c 1 ./ticks
Ticks per usec: 360.727
$ taskset -c 1 ./ticks
Ticks per usec: 345.043

正如我们所见，程序执行之间的差异最多可达 3 倍 (125-360)。这种不稳定性不适用于任何测量。

代码如下(gcc 4.9.3，运行在 Oracle Linux 6.6，内核 3.8.13-55.1.2.el6uek.x86_64):

// g++ -O3 -std=c++11 -Wall ticks.cpp -o ticks
#include <x86intrin.h>
#include <ctime>
#include <cstdint>
#include <iostream>

int main()
{       
    timespec start, end;
    uint64_t s = 0;

    const double rdtsc_ticks_per_usec = [&]()
    {
        unsigned int dummy;

        clock_gettime(CLOCK_MONOTONIC, &start);

        uint64_t rd_start = __rdtscp(&dummy);
        for (size_t i = 0; i < 1000000; ++i) ++s;
        uint64_t rd_end = __rdtscp(&dummy);

        clock_gettime(CLOCK_MONOTONIC, &end);

        double usec_dur = double(end.tv_sec) * 1E6 + end.tv_nsec / 1E3;
        usec_dur -= double(start.tv_sec) * 1E6 + start.tv_nsec / 1E3;

        return (double)(rd_end - rd_start) / usec_dur;
    }();

    std::cout << s << std::endl;
    std::cout << "Ticks per usec: " << rdtsc_ticks_per_usec << std::endl;
    return 0;
}

当我在 Windows 7、i7-4470、VS2015 下运行非常相似的程序时，校准结果非常稳定，只有最后一位的差异很小。

所以问题 - 这个问题是关于什么的？是 CPU 问题、Linux 问题还是我的代码问题？

最佳答案

如果您不确保 cpu 是隔离的，那么还会有其他抖动来源。您确实希望避免在该核心上安排另一个进程。同样理想的是，您运行一个无滴答内核，这样您就永远不会在该内核上运行内核代码。在上面的代码中，我想只有当你不幸在调用 clock_gettime() 和 __rdtscp 之间进行滴答或上下文切换时，这才是重要的

使 s 易变是另一种打败这种编译器优化的方法。

关于c++ - __rdtscp calibration unstable under Linux on Intel Xeon X5550，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/36101311/

24

4

0

文章推荐： javascript - Visual Studio 中 CSS/javascript 的 Vim 样式折叠

文章推荐： css - 删除 Chrome 中的数据列表下拉箭头

文章推荐： python - CSS 解析器 + XHTML 生成器，需要建议

文章推荐： css - Javascript Widget 影响它嵌入的页面的样式

intel-pin - intel pin工具中图像的含义
我是Intel pin工具的新手，最近开始研究pin工具。在教程中，描述了pin工具的模式: Sometimes, however, it can be useful to look at diffe
intel-pin - intel pin工具中图像的含义
我是Intel pin工具的新手，最近开始研究pin工具。在教程中，描述了pin工具的模式: Sometimes, however, it can be useful to look at diffe
intel - 如何开始使用库 intel ipp？
我得到了这份工作:1。产生一个正弦信号。2。使用 FFT 构建其频谱。首先，我为 visual studio 2010 安装了 Intel Parallel Studio XE 2011。在 vs 2
opencl - intel-compute-runtime、intel-opencl-runtime 和 intel-opencl-sdk 之间有什么区别？
看起来 Intel 提供了许多 OpenCL 实现。 ArchWiki描述 OpenCL 实现。它说 beignet 和 intel-opencl 已弃用。那么，intel-compute-runti
intel - 如何读取 "Intel Intrinsics Guide"？
我正在尝试通过阅读 Intel Intrinsics Guide 来开始使用 AVX512 内在函数但到目前为止我发现它没有定义命名数据类型或用于解释的伪代码语法。没有这样的定义，所谓的指南对我起码没
intel - AMD 与 Intel 处理器制作可执行文件
关闭。这个问题是opinion-based 。目前不接受答案。想要改进这个问题吗？更新问题，以便 editing this post 可以用事实和引文来回答它。 . 已关闭 4 年前。 Improv
android-studio - "Intel Atom Image"、 "Google APIs Intel Atom image"和 "Google play Intel Atom Image"之间有什么区别？
在 Android SDK 管理器中，我可以看到 3 种类型的 Intel Atom 图像。有人可以解释“Intel Atom Image”、“Google APIs Intel Atom Image
intel-pin - 使用 intel pintool 记录所有指令
我写了这个 pintool: #include "pin.H" #include #include VOID Instruction(INS ins, VOID *v) { cou
intel - 了解 Intel Intrinsics Guide 中的代码示例
我正在尝试了解 _mm256_permute2f128_ps() 的作用，但无法完全理解 intel's code-example . DEFINE SELECT4(src1, src2, contr
intel - 使用 Intel 内在函数 SSSE3 的替代方案时性能下降
我正在开发一个性能关键应用程序，该应用程序必须移植到仅支持 MMX、SSE、SSE2 和 SSE3 的英特尔凌动处理器中。我以前的应用程序支持 SSSE3 和 AVX，现在我想将其降级为 Intel
intel-pin - Intel Pin 3.0无法识别MPX指令？
我有最新版本的 Intel Pin 3.0 版本 76887。我有一个支持 MPX 的玩具示例: #include int g[10]; int main(int argc, char **arg
intel - 在 Intel 上使用 OpenSolaris 研究 SPARC 可执行结构
我想研究和比较elf、SPARC和PA-RISC的可执行文件结构。为了进行研究，我想在 Intel 机器 (Core2Duo) 上安装 OpenSolaris。但我有一个基本的疑问，它会起作用吗？
intel-mkl - 无法使用 g++ 将数学库与 intel mkl 链接
我尝试使用 g++ 用 intel mkl 11.1 进行编译: g++ -m32 test.c -lmkl_intel -lmkl_intel_thread -lmkl_core -liomp5 -
c++ - 我如何使用 intel 编译器和 intel mpi 安装 boost？
我正在按照以下说明进行操作: https://software.intel.com/en-us/articles/building-boost-with-intel-c-compiler-150 Co
c++ - -masm=intel 标志不适用于使用 Intel 语法在 gcc 编译器中运行汇编语言
我正在尝试在我的 C 程序中使用内联汇编程序 __asm，使用 Intel 语法而不是 AT&T 语法。我正在使用 gcc -S -masm=intel test.c 进行编译但它给出了错误。下面是我
c++ - Intel HD GPU 与 Intel CPU 性能比较
我是 OpenCL 的新手，目前对其性能有一些疑问。我有 Intel(R) Core(TM) i5-4460 CPU @ 3.20GHz + ubuntu + Beignet(Intel 开源 op
Makefile:Intel fortran，文件夹中的源文件，和 Intel Math Kernel Library
我在/ex 文件夹中有一个 main.f90。 f77 子程序文件在/ex/src 中。子程序文件再次使用 BLAS 和 LAPACK 库。对于 BLAS 和 LAPACK，我必须使用英特尔数学核心函
c++ - 为什么此代码链接到 Intel Compiler 2015 而不是 Intel Compiler 2018？
我的团队最近从 2015 年英特尔编译器(并行工作室)升级到 2018 年版本，我们遇到了一个链接器问题，让每个人都焦头烂额。我有以下类(为简洁起见进行了适度编辑)，用于处理子进程的包装以及与它们对
intel - 为什么 Intel Haswell XEON CPU 偶尔会错误计算 FFT 和 ART？
在最后几天，我观察到我无法解释的新工作站的行为。对这个问题做一些研究，INTEL Haswell architecture 中可能存在一个可能的错误。以及在当前的 Skylake Generation
android-emulator - Intel HAXM 安装错误 - 此计算机不支持 Intel 虚拟化技术 (VT-x)
我的 HAXM 安装存在问题。事情是这样的。每次尝试为我的计算机安装 HAXM 时，我都会收到此错误: 问题是，我的计算机支持虚拟化技术(见下图)。知道如何解决这个问题吗？最佳答案只需执行以下步骤

首页

博学

6Ren·AI

商城

c++ - __rdtscp calibration unstable under Linux on Intel Xeon X5550