Intel

Using Intel.com Search

You can easily search the entire Intel.com site in several ways.

  • Brand Name: Core i9
  • Document Number: 123456
  • Code Name: Emerald Rapids
  • Special Operators: “Ice Lake”, Ice AND Lake, Ice OR Lake, Ice*

Quick Links

You can also try the quick links below to see results for most popular searches.

The browser version you are using is not recommended for this site.
Please consider upgrading to the latest version of your browser by clicking one of the following links.

Intel® VTune™ Profiler

ID 766319
Date 5/20/2026
Version
Public
Document Table of Contents
ALU0 Active ALU0 Instructions ALU1 Active ALU1 Instructions ALU2 Active ALU2 Instructions ALU0 and ALU1 Active ALU0 and ALU2 Active ALU0 and XMX Utilization Average Time Computing Threads Started Computing Threads Started, Threads/sec CPU Time EU 2 FPU Pipelines Active EU Array Active EU Array Idle EU Array Stalled/Idle EU Array Stalled EU IPC Rate EU Send pipeline active EU Threads Occupancy Global GPU EU Array Usage GPU HW Scheduler Time GPU Instruction Cache L3 Miss Ratio GPU L3 Atomics GPU L3 Bound GPU L3 Miss Ratio GPU L3 Misses GPU L3 Misses, Misses/sec GPU Load Store Cache Miss Ratio GPU Load Store Cache L3 Miss Ratio GPU LSC Atomics GPU LSC Fences GPU Media Read Requests GPU Media Write Requests GPU SLM Atomics GPU SLM Fences GPU Memory Read Bandwidth, GB/sec GPU Memory Texture Read Bandwidth, GB/sec GPU Memory Write Bandwidth, GB/sec GPU Sampler L3 Miss Ratio GPU Texel Quads Count, Count/sec GPU Utilization Graphics Security Controller Busy Host to GPU Memory Read Bandwidth Host-to-GPU Memory Write Bandwidth Instance Count Instruction Cache Miss Ratio L3 Busy L3 Input Available L3 Instruction Cache Bandwidth L3 Load Store Cache Read Bandwidth L3 Load Store Cache Write Bandwidth L3 Miss Ratio L3 Output Ready L3 Read Bandwidth L3 SQ Full L3 Stalled L3 Write Bandwidth L3 Sampler Bandwidth, GB/sec L3 Shader Bandwidth, GB/sec LLC Miss Rate due GPU Lookups LLC Miss Ratio due GPU Lookups LSC Input Available LSC Output Ready LSC Partial Writes Local Maximum GPU Utilization Multiple Pipe Utilization Occupancy PS EU Active % PS EU Stall % Ratio to Max Bandwidth, % Ratio to Max Bandwidth, % Ratio to Max Bandwidth, % Render/GPGPU Command Streamer Loaded Sampler Input Available Sampler Output Ready Samples Blended Samples Killed in PS, pixels Samples Written Sampler Busy Sampler Is Bottleneck Shared Local Memory Read Bandwidth, GB/sec Shared Local Memory Write Bandwidth, GB/sec SIMD Width SLM Bank Conflicts Stack-to-stack Incoming Bandwidth Stack-to-stack Outgoing Bandwidth System Memory Read Bandwidth System Memory Write Bandwidth Size Total, GB/sec Thread Dispatcher Active TLB Misses Total Time Typed Memory Read Bandwidth, GB/sec Typed Memory Write Bandwidth, GB/sec Typed Reads Coalescence Typed Writes Coalescence Untyped Memory Read Bandwidth, GB/sec Untyped Memory Write Bandwidth, GB/sec Untyped Reads Coalescence Untyped Writes Coalescence Video Codec Busy Video Codec Read Requests Video Codec Write Requests Video Codec 2 Busy Video Codec 2 Read Requests Video Codec 2 Write Requests Video Enhancement Busy Video Enhancement Read Requests Video Enhancement Write Requests Video Enhancement 2 Busy Video Enhancement 2 Read Requests Video Enhancement 2 Write Requests VS EU Active VS EU Stall XVE Barrier Stall XVE Bit Manipulation Instructions XVE Control Stall XVE Dist or Acc Stall XVE INT16\INT32\INT64\FP16\FP32\FP64 Instructions XVE FP16\BF16\INT8\INT4\INT2 XMX Instructions XVE Instruction Fetch Stall XVE Pipe Stall XVE Send Stall XVE SBID Stall XVE XMX Instructions XVE XMX Pipeline Active

Overview

This document explains how you use Intel VTune Profiler to profile serial and multithreaded applications on CPU, GPU, and FPGA platforms. You can run Intel VTune Profiler to profile software applications (local or remote collections) on Windows* and Linux* platforms. Analyze your choice of algorithm and locate or determine:

  • The most time-consuming (hot) functions in your application and/or on the whole system

  • Sections of code that do not effectively utilize available processor time

  • The best sections of code to optimize for sequential performance and for threaded performance

  • Synchronization objects that affect the application performance

  • Whether, where, and why your application spends time on input/output operations

  • Whether your application is CPU or GPU bound and how effectively it offloads code to the GPU

  • The performance impact of different synchronization methods, different numbers of threads, or different algorithms

  • Thread activity and transitions

  • Hardware-related issues in your code such as data sharing, cache misses, branch misprediction, and others

Get Help

Use these resources to get technical support with Intel® VTune™ Profiler and the Intel® oneAPI Toolkit.

Resource

Description

Release Notes

Get information on a specific version of Intel VTune Profiler.

System Requirements

Learn about hardware and software requirements for your use of Intel VTune Profiler.

Analyzers Forum

Discuss your experience with Intel VTune Profiler in the Analyzers Developer Forum. Contact Intel engineers and other users to get questions answered or learn about the product.

Intel VTune Profiler Product Page

Get an overview of Intel VTune Profiler. You can also:

  • Download the product
  • Learn about features in the newest version.
  • Find relevant technical content about workflows and use cases

Intel® oneAPI Toolkit Product Page

Download and get started with the Intel® oneAPI Toolkit.

Developer Resource Page

Use this resource page to get help with all Intel software products.

Tuning Guides and Performance Analysis Papers

Find a collection of tuning guides for use with different Intel microarchitectures.

Read the original on intel.com ↗