Loading... please wait.

Copyright 2011-2026 The Khronos Group Inc.

This Specification is protected by copyright laws and contains material proprietary to Khronos. Except as described by these terms, it or any components may not be reproduced, republished, distributed, transmitted, displayed, broadcast or otherwise exploited in any manner without the express prior written permission of Khronos.

Khronos grants a conditional copyright license to use and reproduce the unmodified Specification for any purpose, without fee or royalty, EXCEPT no licenses to any patent, trademark or other intellectual property rights are granted under these terms.

Khronos makes no, and expressly disclaims any, representations or warranties, express or implied, regarding this Specification, including, without limitation: merchantability, fitness for a particular purpose, non-infringement of any intellectual property, correctness, accuracy, completeness, timeliness, and reliability. Under no circumstances will Khronos, or any of its Promoters, Contributors or Members, or their respective partners, officers, directors, employees, agents or representatives be liable for any damages, whether direct, indirect, special or consequential damages for lost revenues, lost profits, or otherwise, arising from or in connection with these materials.

This Specification has been created under the Khronos Intellectual Property Rights Policy, which is Attachment A of the Khronos Group Membership Agreement available at https://www.khronos.org/files/member_agreement.pdf. Parties desiring to implement the Specification and make use of Khronos trademarks in relation to that implementation, and receive reciprocal patent license protection under the Khronos Intellectual Property Rights Policy must become Adopters and confirm the implementation as conformant under the process defined by Khronos for this Specification; see https://www.khronos.org/adopters.

The Khronos Intellectual Property Rights Policy defines the terms 'Scope', 'Compliant Portion', and 'Necessary Patent Claims'.

Some parts of this Specification are purely informative and so are EXCLUDED from the Scope of this Specification. Section 3.4 defines how these parts of the Specification are identified.

Where this Specification uses technical terminology, defined in the Glossary or otherwise, that refer to enabling technologies that are not expressly set forth in this Specification, those enabling technologies are EXCLUDED from the Scope of this Specification. For clarity, enabling technologies not disclosed with particularity in this Specification (e.g. semiconductor manufacturing technology, hardware architecture, processor architecture or microarchitecture, memory architecture, compiler technology, object oriented technology, basic operating system technology, compression technology, algorithms, and so on) are NOT to be considered expressly set forth; only those application program interfaces and data structures disclosed with particularity are included in the Scope of this Specification.

For purposes of the Khronos Intellectual Property Rights Policy as it relates to the definition of Necessary Patent Claims, all recommended or optional features, behaviors and functionality set forth in this Specification, if implemented, are considered to be included as Compliant Portions.

Where this Specification identifies specific sections of external references, only those specifically identified sections define normative functionality. The Khronos Intellectual Property Rights Policy excludes external references to materials and associated enabling technology not created by Khronos from the Scope of this Specification, and any licenses that may be required to implement such referenced materials and associated technologies must be obtained separately and may involve royalty payments.

Khronos® and Vulkan® are registered trademarks, and 3D Commerce™, ANARI™, Kamaros™, KTX™, glTF™, NNEF™, OpenVG™, OpenVX™, SPIR™, SPIR-V™, SYCL™, Vulkan SC™, and WebGL™ are trademarks of The Khronos Group Inc. OpenXR™ is a trademark owned by The Khronos Group Inc. and is registered as a trademark in China, the European Union, Japan and the United Kingdom. OpenCL™ is a trademark of Apple Inc. used under license by Khronos. OpenGL® is a registered trademark and the OpenGL ES™ and OpenGL SC™ logos are trademarks of Hewlett Packard Enterprise used under license by Khronos. ASTC is a trademark of ARM Holdings PLC. All other product names, trademarks, and/or company names are used solely for identification and belong to their respective owners.

1. Acknowledgements

Editors

  • Maria Rovatsou, Codeplay

  • Lee Howes, Qualcomm

  • Ronan Keryell, AMD

  • Greg Lueck, Intel (current)

Contributors

  • Eric Berdahl, Adobe

  • Shivani Gupta, Adobe

  • David Neto, Altera

  • Carlo Bertolli, AMD

  • Andrew Gozillon, AMD

  • Gauthier Harnisch, AMD

  • Ronan Keryell, AMD

  • Yiannis Papadopoulos, AMD

  • Brian Sumner, AMD

  • Lin-Ya Yu, AMD

  • Thomas Applencourt, Argonne National Laboratory

  • Hal Finkel, Argonne National Laboratory

  • Kevin Harms, Argonne National Laboratory

  • Michael Lance, Argonne National Laboratory

  • Nevin Liber, Argonne National Laboratory

  • Anastasia Stulova, ARM

  • Balázs Keszthelyi, Broadcom

  • Alexandra Crabb, Caster Communications

  • Aymeric Millan, CEA, Maison de la Simulation

  • Stuart Adams, Codeplay

  • Verena Beckham, Codeplay, Qualcomm

  • Aidan Belton, Codeplay

  • Gordon Brown, Codeplay

  • Hugh Delaney, Codeplay

  • Atharva Dubey, Codeplay

  • Morris Hafner, Codeplay

  • Alexander Johnston, Codeplay

  • Marios Katsigiannis, Codeplay

  • Paul Keir, Codeplay

  • Steffen Larsen, Codeplay

  • Victor Lomüller, Codeplay

  • Tomas Matheson, Codeplay

  • Duncan McBain, Codeplay

  • Nicolas Miller, Codeplay

  • Georgi Mirazchiyski, Codeplay

  • Mahmoud Moadeli, Codeplay

  • Ralph Potter, Codeplay

  • Ruyman Reyes, Codeplay

  • Andrew Richards, Codeplay

  • Maria Rovatsou, Codeplay

  • Panagiotis Stratis, Codeplay

  • Erik Tomusk, Codeplay

  • Michael Wong, Codeplay

  • Peter Žužek, Codeplay

  • Matt Newport, EA

  • Rasool Maghareh, Huawei Technologies Co. Ltd.

  • Guansong Zhang, Huawei Technologies Co. Ltd.

  • Ruslan Arutyunyan, Intel

  • Alexey Bader, Intel

  • James Brodman, Intel

  • Ilya Burylov, Intel

  • Jessica Davies, Intel

  • Andrei Elovikov, Intel

  • Felipe de Azevedo Piovezan, Intel

  • Allen Hux, Intel

  • Michael Kinsner, Intel

  • Nikita Kornev, Intel

  • Greg Lueck, Intel

  • John Pennycook, Intel

  • Pablo Reble, Intel

  • Roland Schulz, Intel

  • Sergey Semenov, Intel

  • Jason Sewall, Intel

  • James O’Riordon, Khronos

  • Jon Leech, Luna Princeps LLC

  • Kathleen Mattson, Miller & Mattson, LLC

  • Dave Miller, Miller & Mattson, LLC

  • Stéphanie Even, Mercedes-Benz Research and Development NA

  • Chris Gearing, Mobileye

  • Seiji Nishimura, NSITEXE, Inc.

  • Neil Trevett, NVIDIA

  • Lee Howes, Qualcomm

  • Chu-Cheow Lim, Qualcomm

  • Jack Liu, Qualcomm

  • Hongqiang Wang, Qualcomm

  • Ruihao Zhang, Qualcomm

  • Dave Airlie, Red Hat

  • Hyesun Hong, Samsung Electronics

  • Aksel Alpay, Self

  • Dániel Berényi, Self

  • Nuno Nobre, STFC Hartree Centre

  • Máté Nagy-Egri, Stream HPC

  • Bálint Soproni, Stream HPC

  • Tom Deakin, University of Bristol

  • Philip Salzmann, University of Innsbruck

  • Peter Thoman, University of Innsbruck

  • Biagio Cosenza, University of Salerno

  • Paul Preney, University of Windsor

2. Introduction

SYCL (pronounced “sickle”) is a royalty-free, cross-platform API for heterogeneous computing in C++.

SYCL enables developers to write standard C++ code that executes on a wide range of devices, using modern techniques such as inheritance, templates, and lambda expressions. All computational kernels to be executed on a device can be written inside C++ source files as normal C++ code, alongside any code intended to be run on a system’s host processor. This concept, known as “single-source” programming, reduces the complexity of heterogeneous programming for developers and gives compilers greater opportunities to analyze/optimize across the host-device boundary.

SYCL is designed to be as close to standard C++ as possible, and some implementations of SYCL may be able to use a standard C++ compiler to target CPU devices. However, to ensure portability of device code across a wide range of devices, SYCL imposes some restrictions on the set of C++ features that SYCL implementations are required to support within device code. These restrictions may not be applicable to all devices and can therefore be relaxed by specific Khronos extensions or vendor extensions.

SYCL was originally based on OpenCL, and retains an execution model, runtime feature set, and device capability set inspired by the OpenCL standard. However, there is no requirement that SYCL implementations must use OpenCL; SYCL implementations are free to support devices via any low-level API (or “backend”) they choose.

Some of the key features of SYCL are:

  • Common parallel patterns, such as reductions and group algorithms, are exposed via high-level abstractions.

  • Interoperability with the lower-level capabilities of specific backends guarantees access to platform-specific optimizations.

  • Buffers and accessors provide a simple way to build task-graphs without manually managing dependencies.

  • Unified Shared Memory (USM) provides an explicit, pointer-based, mechanism for managing and sharing data.

SYCL has been designed to enable implementations on a wide variety of platforms, permitting easy integration with other platform-specific technologies. Both users and implementers are encouraged to build upon SYCL as an open platform for system-wide heterogeneous programming.

3. SYCL architecture

This chapter describes the structure of a SYCL application, and how the SYCL generic programming model lays out on top of a number of SYCL backends.

3.1. Overview

SYCL is an open industry standard for programming a heterogeneous system. The design of SYCL allows standard C++ source code to be written such that it can run on either an heterogeneous device or on the host.

The terminology used for SYCL inherits historically from OpenCL with some SYCL-specific additions. However SYCL is a generic C++ programming model that can be laid out on top of other APIs apart from OpenCL. SYCL implementations can provide SYCL backends for various APIs, implementing the SYCL general specification on top of them. We refer to this API as the SYCL backend API. The SYCL general specification defines the behavior that all SYCL implementations must expose to SYCL users for a SYCL application to behave as expected.

A function object that can execute on a device exposed by a SYCL backend API is called a SYCL kernel function.

To ensure maximum interoperability with different SYCL backend APIs, software developers can access the SYCL backend API alongside the SYCL general API whenever they include the SYCL backend interoperability headers. However, interoperability is a SYCL backend-specific feature. An application that uses interoperability does not conform to the SYCL general application model, since it is not portable across backends.

The target users of SYCL are C++ programmers who want all the performance and portability features of a standard like OpenCL, but with the flexibility to use higher-level C++ abstractions across the host/device code boundary. Developers can use most of the abstraction features of C++, such as templates, classes and operator overloading.

However, some C++ language features are not permitted inside kernels, due to the limitations imposed by the capabilities of the underlying heterogeneous platforms. These features include virtual functions, virtual inheritance, throwing/catching exceptions, and run-time type-information. These features are available outside kernels as normal. Within these constraints, developers can use abstractions defined by SYCL, or they can develop their own on top. These capabilities make SYCL ideal for library developers, middleware providers and application developers who want to separate low-level highly-tuned algorithms or data structures that work on heterogeneous systems from higher-level software development. Software developers can produce templated algorithms that are easily usable by developers in other fields.

3.2. Anatomy of a SYCL application

Below is an example of a typical SYCL application which schedules a job to run in parallel on any device available.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
#include <iostream>
#include <sycl/sycl.hpp>
#include <vector>

int main() {
  // Declare number of work items
  constexpr size_t N = 1024;

  // Allocate host memory to store the results
  std::vector<int> dataHost(N);

  // Create an in order queue to enqueue work to the default device
  sycl::queue myQueue{sycl::property::queue::in_order()};

  // Allocate device memory to be worked on
  int *dataDevice = sycl::malloc_device<int>(N, myQueue);

  // Enqueue a parallel_for task with 1024 work-items
  myQueue.parallel_for(N, [=](sycl::id<1> idx) {
    // Initialize each buffer element with its own rank number starting at 0
    dataDevice[idx] = idx;
  }); // End of the kernel function

  // Copy the results back to the host from the device
  myQueue.copy(dataDevice, dataHost.data(), N);

  myQueue.wait(); // Wait for the queue to finish executing all the tasks

  // Print result
  for (int i = 0; i < N; i++)
    std::cout << "dataHost[" << i << "] = " << dataHost[i] << std::endl;

  // Free device memory
  sycl::free(dataDevice, myQueue);

  return 0;
}

At line 2 we #include the SYCL header files, which provide all of of the SYCL features that will be used.

At line 13 we instantiate a queue with the in_order property. A queue is bound to a device, if you do not specify a device with a selector function, the SYCL runtime will choose one automatically. A queue will execute commands submitted to it on it’s associated device. In our case commands will be executed in a first in first out (FIFO) fashion because of the in_order property passed to our queue’s constructor. The commands submitted to myQueue are executed asynchronously from the host code.

At line 16 we use malloc_device to allocate memory on the device associated with our queue and store a unified shared memory (USM) pointer to the memory in dataDevice. This memory is only accessible on the device.

In lines 19 to 22, parallel_for does three conceptual things:

  1. The SYCL kernel function (defined here as a lambda), passed as the second argument to parallel_for, is compiled by a device compiler into a kernel that can be run on the device associated with myQueue.

  2. A command that invokes the kernel on the device is submitted to myQueue, and the kernel is executed asynchronously from the host.

  3. The command decomposes the range 0 to 1023 into work-items. For each work-item the SYCL kernel function is invoked once on the device, and these invocations may be executed in parallel.

The different queue member functions used to invoke a SYCL kernel function can be found in Section 4.9.4.2.

At line 25 we use the copy member function to submit a copy command to copy the results from the device’s memory back into host memory.

At line 27 we call wait to block the host until all commands submitted to myQueue (the parallel_for and the copy) have completed.

In lines 30 and 31 we print the results from the host memory, which now contains a copy of the device memory.

At line 34 the device allocation dataDevice is released with free.

3.3. Normative references

The documents in the following list are referred to within this SYCL specification, and their content is a requirement for this document.

  1. C++17: ISO/IEC 14882:2017 Clauses 1-19, referred to in this specification as the C++ core language. The SYCL specification refers to language in the following C++ defect reports and assumes a compiler that implements them: DR2325.

  2. C++20: ISO/IEC 14882:2020 Programming languages — C++, referred to in this specification as the next C++ specification.

3.4. Non-normative notes and examples

Unless stated otherwise, text within this SYCL specification is normative and defines the required behavior of a SYCL implementation. Non-normative / informational notes are included within this specification using either of two formats. One format for non-normative notes is the “note” callout of this form:

Information within a note callout, such as this text, is for informational purposes and does not impose requirements on or specify behavior of a SYCL implementation.

The other format for a non-normative note is like this:

[Note: This is also a non-normative note. — end note]

Source code examples within the specification are provided to aid with understanding, and are non-normative.

In case of any conflict between a non-normative note or source example, and normative text within the specification, the normative text must be taken to be correct.

3.5. The SYCL platform model

The SYCL platform model consists of a host connected to one or more devices, called devices. Devices are grouped together into one or multiple platforms. An implementation may also expose empty platforms that do not contain any devices.

A SYCL context is constructed, either directly by the user or implicitly when creating a queue, to hold all the runtime information required by the SYCL runtime and the SYCL backend to operate on a device, or group of devices. When a group of devices can be grouped together on the same context, they have some visibility of each other’s memory objects. The SYCL runtime can assume that memory is visible across all devices in the same context.

A SYCL application executes on the host as a standard C++ program. Devices are exposed through different SYCL backends to the SYCL application. The SYCL application submits command group function objects to queues. Each queue enables execution on a given device.

The SYCL runtime then extracts operations from the command group function object, e.g. an explicit copy operation or a SYCL kernel function. When the operation is a SYCL kernel function, the SYCL runtime uses a SYCL backend-specific mechanism to extract the device binary from the SYCL application and pass it to the SYCL backend API for execution on the device.

A SYCL device is divided into one or more compute units (CUs) which are each divided into one or more processing elements (PEs). Computations on a device occur within the processing elements. How computation is mapped to PEs is SYCL backend and device specific. Two devices exposed via two different backends can map computations differently to the same device.

When a SYCL application contains SYCL kernel function objects, the SYCL implementation must provide an offline compilation mechanism that enables the integration of the device binaries into the SYCL application. The output of the offline compiler can be an intermediate representation, such as SPIR-V, that will be finalized during execution or a final device ISA.

A device may expose special purpose functionality as a built-in function. The SYCL API exposes functions to query and dispatch said built-in functions. Some SYCL backends and devices may not support programmable kernels, and only support built-in functions.

3.6. The SYCL backend model

SYCL is a generic programming model for the C++ language that can target multiple APIs, such as OpenCL.

SYCL implementations enable these target APIs by implementing SYCL backends. For a SYCL implementation to be conformant on said SYCL backend, it must execute the SYCL generic programming model on the backend. All SYCL implementations must provide at least one backend.

The present document covers the SYCL generic interface available to all SYCL backends. How the SYCL generic interface maps to a particular SYCL backend is defined either by a separate SYCL backend specification document, provided by the Khronos SYCL group, or by the SYCL implementation documentation. Whenever there is a SYCL backend specification document, this takes precedence over SYCL implementation documentation.

When a SYCL user builds their SYCL application, she decides which of the SYCL backends will be used to build the SYCL application. This is called the set of active backends. Implementations must ensure that the active backends selected by the user can be used simultaneously by the SYCL implementation at runtime. If two backends are available at compile time but will produce an invalid SYCL application at runtime, the SYCL implementation must emit a compilation error.

A SYCL application built with a number of active backends does not necessarily guarantee that said backends can be executed at runtime. The subset of active backends available at runtime is called available backends. A backend is said to be available if the host platform where the SYCL application is executed exposes support for the API required for the SYCL backend.

It is implementation dependent whether certain backends require third-party libraries to be available in the system. Failure to have all dependencies required for all active backends at runtime will cause the SYCL application to not run.

Once the application is running, users can query what SYCL platforms are available. SYCL implementations will expose the devices provided by each backend grouped into platforms. A backend must expose at least one platform.

Under the SYCL backend model, SYCL objects can contain one or multiple references to a certain SYCL backend native type. Not all SYCL objects will map directly to a SYCL backend native type. The mapping of SYCL objects to SYCL backend native types is defined by the SYCL backend specification document when available, or by the SYCL implementation otherwise.

To guarantee that multiple SYCL backend objects can interoperate with each other, SYCL memory objects are not bound to a particular SYCL backend. SYCL memory objects can be accessed from any device exposed by an available backend. SYCL Implementations can potentially map SYCL memory objects to multiple native types in different SYCL backends.

Since SYCL memory objects are independent of any particular SYCL backend, SYCL command groups can request access to memory objects allocated by any SYCL backend, and execute it on the backend associated with the queue. This requires the SYCL implementation to be able to transfer memory objects across SYCL backends.

USM allocations are subject to the limitations described in Section 4.8.

When a SYCL application runs on any number of SYCL backends without relying on any SYCL backend-specific behavior or interoperability, it is said to be a SYCL general application, and it is expected to run in any SYCL-conformant implementation that supports the required features for the application.

3.6.1. Platform mixed version support

The SYCL generic programming model exposes a number of platforms, each of them either empty or exposing a number of devices. Each platform is bound to a certain SYCL backend. SYCL devices associated with said platform are associated with that SYCL backend.

Although the APIs in the SYCL generic programming model are defined according to this specification and their version is indicated by the macro SYCL_LANGUAGE_VERSION, this does not apply to APIs exposed by the SYCL backends. Each SYCL backend provides its own document that defines its APIs, and that document tells how to query for the device and platform versions.

3.7. SYCL execution model

As described in Section 3.2, a SYCL application is comprised of three scopes: application scope, command group scope, and kernel scope. Code in the application scope and command group scope runs on the host and is governed by the SYCL application execution model. Code in the kernel scope runs on a device and is governed by the SYCL kernel execution model.

A SYCL device does not necessarily correspond to a physical accelerator. A SYCL implementation may choose to expose some or all of the host’s resources as a SYCL device; such an implementation would execute code in kernel scope on the host, but that code would still be governed by the SYCL kernel execution model.

3.7.1. SYCL application execution model

The SYCL application defines the execution order of the kernels by grouping each kernel with its requirements into a command group function object. Command group function objects are submitted for execution via a queue object, which defines the device where the kernel will run. This specification sometimes refers to this as “submitting the kernel to a device”. The same command group object can be submitted to different queues. When a command group is submitted to a SYCL queue, the requirements of the kernel execution are captured. The implementation can start executing a kernel as soon as its requirements have been satisfied.

3.7.1.1. Backend resources managed by the SYCL application

The SYCL runtime integrated with the SYCL application will manage the resources required by the SYCL backend API to manage the devices it is providing access to. This includes, but is not limited to, resource handlers, memory pools, dispatch queues and other temporary handler objects.

The SYCL programming interface represents the lifetime of the resources managed by the SYCL application using RAII rules. Construction of a SYCL object will typically entail the creation of multiple SYCL backend objects, which will be properly released on destruction of said SYCL object. The overall rules for construction and destruction are detailed in Chapter 4. Those SYCL backends with a SYCL backend document will detail how the resource management from SYCL objects map down to the SYCL backend objects.

In SYCL, the minimum required object for submitting work to devices is the queue, which contains references to a platform, device and a context internally.

The resources managed by SYCL are:

  1. Platforms: all features of SYCL backend APIs are implemented by platforms. A platform can be viewed as a given vendor’s runtime and the devices accessible through it. Some devices will only be accessible to one vendor’s runtime and hence multiple platforms may be present. SYCL manages the different platforms for the user which are accessible through a sycl::platform object. In some cases, an implementation might also choose to expose empty sycl::platform objects, for example if a vendor’s runtime is available, but no devices supported by that runtime are available in the system.

  2. Contexts: any SYCL backend resource that is acquired by the user is attached to a context. A context contains a collection of devices that the host can use and manages memory objects that can be shared between the devices. Devices belonging to the same context must be able to access each other’s global memory using some implementation-specific mechanism. A given context can only wrap devices owned by a single platform. A context is exposed to the user with a sycl::context object.

  3. Devices: platforms may provide devices for executing SYCL kernels. In SYCL, a device is accessible through a sycl::device object.

  4. Kernels: the SYCL functions that run on SYCL devices are defined as C++ function objects (a named function object type or a lambda expression). A kernel can be introspected through a sycl::kernel object.

    Note that some SYCL backends may expose non-programmable functionality as pre-defined kernels.

  5. Kernel bundles: Kernels are stored internally in the SYCL application as device images, and these device images can be grouped into a sycl::kernel_bundle object. These objects provide a way for the application to control the online compilation of kernels for devices.

  6. Queues: SYCL kernels execute in command queues. The user must create a sycl::queue object, which references an associated context, platform and device. The context, platform and device may be chosen automatically, or specified by the user. SYCL queues execute kernels on a particular device of a particular context, but can have dependencies from any device on any available SYCL backend.

The SYCL implementation guarantees the correct initialization and destruction of any resource handled by the underlying SYCL backend API, except for those the user has obtained manually via the SYCL interoperability API.

3.7.1.2. SYCL command groups and execution order

By default, SYCL queues execute kernel functions in an out-of-order fashion based on dependency information. Developers only need to specify what data is required to execute a particular kernel. The SYCL runtime will guarantee that kernels are executed in an order that guarantees correctness. By specifying access modes and types of memory, a directed acyclic dependency graph (DAG) of kernels is built at runtime. This is achieved via the usage of command group objects. A SYCL command group object defines a set of requisites (R) and a kernel function (k). A command group is submitted to a queue when using the sycl::queue::submit member function.

A requisite (ri) is a requirement that must be fulfilled for a kernel-function (k) to be executed on a particular device. For example, a requirement may be that certain data is available on a device, or that another command group has finished execution. An implementation may evaluate the requirements of a command group at any point after it has been submitted. The processing of a command group is the process by which a SYCL runtime evaluates all the requirements in a given R. The SYCL runtime will execute k only when all ri are satisfied (i.e., when all requirements are satisfied). To simplify the notation, in the specification we refer to the set of requirements of a command group named foo as CGfoo = r1, …, rn.

The evaluation of a requisite (Satisfied(ri)) returns the status of the requisite, which can be True or False. A satisfied requisite implies the requirement is met. Satisfied(ri) never alters the requisite, only observes the current status. The implementation may not block to check the requisite, and the same check can be performed multiple times.

An action (ai) is a collection of implementation-defined operations that must be performed in order to satisfy a requisite. The set of actions for a given command group A is permitted to be empty if no operation is required to satisfy the requirement. The notation ai represents the action required to satisfy ri. Actions of different requisites can be satisfied in any order with respect to each other without side effects (i.e., given two requirements rj and rk, (rj, rk)(rk, rj)). The intersection of two actions is not necessarily empty. Actions can include (but are not limited to): memory copy operations, memory mapping operations, coordination with the host, or implementation-specific behavior.

Finally, Performing an action (Perform(ai)) executes the action operations required to satisfy the requisite rj. Note that, after Perform(ai), the evaluation Satisfied(rj) will return True until the kernel is executed. After the kernel execution, it is not defined whether a different command group with the same requirements needs to perform the action again, where actions of different requisites inside the same command group object can be satisfied in any order with respect to each other without side effects: Given two requirements rj and rk, Perform(aj) followed by Perform(ak) is equivalent to Perform(ak) followed by Perform(aj).

The requirements of different command groups submitted to the same or different queues are evaluated in the relative order of submission. command group objects whose intersection of requirement sets is not empty are said to depend on each other. They are executed in order of submission to the queue. If command groups are submitted to different queues or by multiple threads, the order of execution is determined by the SYCL runtime. Note that independent command group objects can be submitted simultaneously without affecting dependencies.

Table 1 illustrates the execution order of three command group objects (CGa,CGb,CGc) with certain requirements submitted to the same queue. Both CGa and CGb only have one requirement, r1 and r2 respectively. CGc requires both r1 and r2. This enables the SYCL runtime to potentially execute CGa and CGb simultaneously, whereas CGc cannot be executed until both CGa and CGb have been completed. The SYCL runtime evaluates the requisites and performs the actions required (if any) for the CGa and CGb. When evaluating the requisites of CGc, they will be satisfied once the CGa and CGb have finished.

Table 1. Execution order of three command groups submitted to the same queue
SYCL Application Enqueue Order SYCL Kernel Execution Order
sycl::queue syclQueue;
syclQueue.submit(CGa(r1));
syclQueue.submit(CGb(r2));
syclQueue.submit(CGc(r1,r2));
three cg one queue

Table 2 uses three separate SYCL queue objects to submit the same command group objects as before. Regardless of using three different queues, the execution order of the different command group objects is the same. When different threads enqueue to different queues, the execution order of the command group will be the order in which the submit member functions are executed. In this case, since the different command group objects execute on different devices, the actions required to satisfy the requirements may be different (e.g, the SYCL runtime may need to copy data to a different device in a separate context).

Table 2. Execution order of three command groups submitted to the different queues
SYCL Application Enqueue Order SYCL Kernel Execution Order
sycl::queue syclQueue1;
sycl::queue syclQueue2;
sycl::queue syclQueue3;
syclQueue1.submit(CGa(r1));
syclQueue2.submit(CGb(r2));
syclQueue3.submit(CGc(r1,r2));
three cg three queue
3.7.1.3. Controlling execution order with events

Submitting an action for execution returns an event object. Programmers may use these events to explicitly coordinate host and device execution. Host code can wait for an event to complete, which will block execution on the host until the action(s) represented by the event have completed. The event class is described in greater detail in Section 4.6.6.

Events may also be used to explicitly order the execution of kernels. Host code may wait for the completion of specific event, which blocks execution on the host until that event’s action has completed. Events may also define requisites between command groups. Using events in this manner informs the runtime that one or more command groups must complete before another command group may begin executing. See Section 4.9.4.1 for greater detail.

3.7.2. SYCL kernel execution model

When a kernel is submitted for execution, an index space is defined. An instance of the kernel body executes for each point in this index space. This kernel instance is called a work-item and is identified by its point in the index space, which provides a global id for the work-item. Each work-item executes the same code but the specific execution pathway through the code and the data operated upon can vary by using the work-item global id to specialize the computation.

An index space of size zero is allowed. All aspects of kernel execution proceed as normal with the exception that the kernel function itself is not executed. Note this means the command queue will still schedule this kernel after satisfying the requirements and this satisfies requirements of any dependent enqueued kernels.

3.7.2.1. Basic kernels

SYCL allows a simple execution model in which a kernel is invoked over an N-dimensional index space defined by range<N>, where N is one, two or three. Each work-item in such a kernel executes independently.

Each work-item is identified by a value of type item<N>. The type item<N> encapsulates a work-item identifier of type id<N> and a range<N> representing the number of work-items executing the kernel.

3.7.2.2. ND-range kernels

Work-items can be organized into work-groups, providing a more coarse-grained decomposition of the index space. Each work-group is assigned a unique work-group id with the same dimensionality as the index space used for the work-items. Work-items are each assigned a local id, unique within the work-group, so that a single work-item can be uniquely identified by its global id or by a combination of its local id and work-group id. The work-items in a given work-group execute on the processing elements of a single compute unit.

When work-groups are used in SYCL, the index space is called an nd-range. An ND-range is an N-dimensional index space, where N is one, two or three. In SYCL, the ND-range is represented via the nd_range<N> class. An nd_range<N> is made up of a global range and a local range, each represented via values of type range<N>. Additionally, there can be a global offset, represented via a value of type id<N>; this is deprecated in SYCL 2020. The types range<N> and id<N> are each N-element arrays of integers. The iteration space defined via an nd_range<N> is an N-dimensional index space starting at the ND-range’s global offset whose size is its global range, split into work-groups of the size of its local range.

Each work-item in the ND-range is identified by a value of type nd_item<N>. The type nd_item<N> encapsulates a global id, local id and work-group id, all of type id<N> (the iteration space offset also of type id<N>, but this is deprecated in SYCL 2020), as well as global and local ranges and coordination mechanisms necessary to make work-groups useful. Work-groups are assigned ids using a similar approach to that used for work-item global ids. Work-items are assigned to a work-group and given a local id with components in the range from zero to the size of the work-group in that dimension minus one. Hence, the combination of a work-group id and the local id within a work-group uniquely defines a work-item.

3.7.2.3. Backend-specific kernels

SYCL allows a SYCL backend to expose fixed functionality as non-programmable built-in kernels. The availability and behavior of these built-in kernels are SYCL backend-specific, and are not required to follow the SYCL execution and memory models. Furthermore the interface exposed utilize these built-in kernels is also SYCL backend-specific. See the relevant backend specification for details.

3.8. Memory model

Since SYCL is a single-source programming model, the memory model affects both the application and the device kernel parts of a program. On the SYCL application, the SYCL runtime will make sure data is available for execution of the kernels. On the SYCL device kernel, the SYCL backend rules describing how the memory behaves on a specific device are mapped to SYCL C++ constructs. Thus it is possible to program kernels efficiently in pure C++.

3.8.1. SYCL application memory model

The application running on the host uses SYCL buffer objects using instances of the sycl::buffer class or USM allocation functions to allocate memory in the global address space, or can allocate specialized image memory using the sycl::unsampled_image and sycl::sampled_image classes.

In the SYCL application, memory objects are bound to all devices in which they are used, regardless of the SYCL context where they reside. SYCL memory objects (namely, buffer and image objects) can encapsulate multiple underlying SYCL backend memory objects together with multiple host memory allocations to enable the same object to be shared between devices in different contexts, platforms or backends. USM allocations uniquely identify a memory allocation and are bound to a SYCL context. They are only valid on the backend used by the context.

The order of execution of command group objects ensures a sequentially consistent access to the memory from the different devices to the memory objects. Accessing a USM allocation does not alter the order of execution. Users must explicitly inform the SYCL runtime of any requirements necessary for a legal execution.

To access a memory object, the user must create an accessor object which parameterizes the type of access to the memory object that a kernel or the host requires. The accessor object defines a requirement to access a memory object, and this requirement is defined by construction of an accessor, regardless of whether there are any uses in a kernel or by the host. An accessor object specifies whether the access is via global memory, constant memory or image samplers and their associated access functions. The accessor also specifies whether the access is read-only (RO), write-only (WO) or read-write (RW). An optional no_init property can be added to an accessor to tell the system to discard any previous contents of the data the accessor refers to, so there are two additional requirement types: no-init-write-only (NWO) and no-init-read-write (NRW). For simplicity, when a requisite represents an accessor object in a certain access mode, we represent it as MemoryObjectAccessMode. For example, an accessor that accesses memory object buf1 in RW mode is represented as buf1RW. A command group object that uses such an accessor is represented as CG(buf1RW). The action required to satisfy a requisite and the location of the latest copy of a memory object will vary depending on the implementation.

Table 3 illustrates an example where command group objects are enqueued to two separate SYCL queues executing in devices in different contexts. The requisites for the command group execution are the same, but the actions to satisfy them are different. For example, if the data is on the host before execution, A(b1RW) and A(b2RW) can potentially be implemented as copy operations from the host memory to context1 or context2 respectively. After CGa and CGb are executed, A'(b1RW) will likely be an empty operation, since the result of the kernel can stay on the device. On the other hand, the results of CGb are now on a different context than CGc is executing, therefore A'(b2RW) will need to copy data across two separate contexts using an implementation specific mechanism.

Table 3. Actions performed when three command groups are submitted to two distinct queues
SYCL Application Enqueue Order SYCL Kernel Execution Order
sycl::queue q1(context1);
sycl::queue q2(context2);
q1.submit(CGa(b1RW));
q2.submit(CGb(b2RW));
q1.submit(CGc(b1RW,b2RW));
device to device1

Possible implementation by a SYCL Runtime

device to device2

Table 3 shows actions performed when three command groups are submitted to two distinct queues, and potential implementation in an OpenCL SYCL backend by a SYCL runtime. Note that in this example, each SYCL buffer (b2,b2) is implemented as separate cl_mem objects per context.

Note that the order of the definition of the accessors within the command group is irrelevant to the requirements they define. All accessors always apply to the entire command group object where they are defined.

When multiple accessors in the same command group define different requisites to the same memory object these requisites must be resolved.

Firstly, any requisites with different access modes but the same access target are resolved into a single requisite with the union of the different access modes according to Table 4. The atomic access mode acts as if it was read-write (RW) when determining the combined requirement. The rules in Table 4 are commutative and associative.

Table 4. Combined requirement from two different accessor access modes within the same command group. The rules are commutative and associative
One access mode Other access mode Combined requirement

read (RO)

write (WO)

read-write (RW)

read (RO)

read-write (RW)

read-write (RW)

write (WO)

read-write (RW)

read-write (RW)

no-init-write (NWO)

no-init-read-write (NRW)

no-init-read-write (NRW)

no-init-write (NWO)

write (WO)

write (WO)

no-init-write (NWO)

read (RO)

read-write (RW)

no-init-write (NWO)

read-write (RW)

read-write (RW)

no-init-read-write (NRW)

write (WO)

read-write (RW)

no-init-read-write (NRW)

read (RO)

read-write (RW)

no-init-read-write (NRW)

read-write (RW)

read-write (RW)

The result of this should be that there should not be any requisites with the same access target.

Secondly, the remaining requisites must adhere to the following rule. Only one of the requisites may have write access (W or RW), otherwise the SYCL runtime must throw an exception. All requisites create a requirement for the data they represent to be made available in the specified access target, however only the requisite with write access determines the side effects of the command group, i.e. only the data which that requisite represents will be updated.

For example:

  • CG(b1GRW, b1HR) is permitted.

  • CG(b1GRW, b1HRW) is not permitted.

  • CG(b1GW, b1CRW) is not permitted.

Where G and C correspond to a target::device and target::constant_buffer accessor and H corresponds to a host accessor.

A buffer created from a range of an existing buffer is called a sub-buffer. A buffer may be overlaid with any number of sub-buffers. Accessors can be created to operate on these sub-buffers. Refer to Section 4.7.2 for details on sub-buffer creation and restrictions. A requirement to access a sub-buffer is represented by specifying its range, e.g. CG(b1RW,[0,5)) represents the requirement of accessing the range [0,5) buffer b1 in read write mode.

If two accessors are constructed to access the same buffer, but both are to non-overlapping sub-buffers of the buffer, then the two accessors are said to not overlap, otherwise the accessors do overlap. Overlapping is the test that is used to determine the scheduling order of command groups. Command-groups with non-overlapping requirements may execute concurrently.

Table 5. Requirements on overlapping vs non-overlapping sub-buffer
SYCL Application Enqueue Order SYCL Kernel Execution Order
sycl::queue q1(context1);
q1.submit(CGa(b1{RW,[0,10)}));
q1.submit(CGb(b1{RW,[10,20)));
q1.submit(CGc(b1RW,[5,15)));
overlap

It is permissible for command groups that only read data to not copy that data back to the host or other devices after reading and for the runtime to maintain multiple read-only copies of the data on multiple devices.

A special case of requirement is the one defined by a host accessor. Host accessors are represented with H(MemoryObjectAccessMode), e.g, H(b1RW) represents a host accessor to b1 in read-write mode. Host accessors are a special type of accessor constructed from a memory object outside a command group, and require that the data associated with the given memory object is available on the host in the given pointer. This causes the runtime to block on construction of this object until the requirement has been satisfied. Host accessor objects are effectively barriers on all accesses to a certain memory object. Table 6 shows an example of multiple command groups enqueued to the same queue. Once the host accessor H(b1RW) is reached, the execution cannot proceed until CGa is finished. However, CGb does not have any requirements on b1, therefore, it can execute concurrently with the barrier. Finally, CGc will be enqueued after H(b1RW) is finished, but still has to wait for CGb to conclude for all its requirements to be satisfied. See Section 3.9.8 for details on host-device coordination.

Table 6. Execution of command groups when using host accessors
SYCL Application Enqueue Order SYCL Kernel Execution Order
sycl::queue q1;
q1.submit(CGa(b1RW));
q1.submit(CGb(b2RW));

H(b1RW);

q1.submit(CGc(b1RW, b2RW));
host acc

3.8.2. SYCL device memory model

The memory model for SYCL devices is based on the OpenCL memory model. Work-items executing in a kernel have access to three distinct address spaces (memory regions) and a virtual address space overlapping some concrete address spaces:

  • Global-memory is accessible to all work-items in all work-groups. Work-items can read from or write to any element of a global memory object. Reads and writes to global memory may be cached depending on the capabilities of the device. Global memory is persistent across kernel invocations. Concurrent access to a location in an USM allocation by two or more executing kernels where at least one kernel modifies that location is a data race; there is no guarantee of correct results unless mem-fence and atomic operations are used.

  • Local-memory is accessible to all work-items in a single work-group. Attempting to access local memory in one work-group from another work-group results in undefined behavior. This memory region can be used to allocate variables that are shared by all work-items in a work-group. Work-group-level visibility allows local memory to be implemented as dedicated regions of the device memory where this is appropriate.

  • Private-memory is a region of memory private to a work-item. Attempting to access private memory in one work-item from another work-item results in undefined behavior.

  • Generic-memory is a virtual address space which overlaps the global, local and private address spaces. Therefore, an object that resides in the global, local, or private address space can also be accessed through the generic address space.

3.8.2.1. Access to memory

Accessors in the device kernels provide access to the memory objects, acting as pointers to the corresponding address space.

Pointers can be passed directly as kernel arguments if an implementation supports USM. See Section 4.8 for information on when it is legal to dereference pointers passed from the host inside kernels.

To allocate local memory within a kernel, the user can either pass a sycl::local_accessor object as a argument to an ND-range kernel (that has a user-defined work-group size), or can define a variable in work-group scope inside sycl::parallel_for_work_group.

Any variable defined inside a sycl::parallel_for scope or sycl::parallel_for_work_item scope will be allocated in private memory. Any variable defined inside a sycl::parallel_for_work_group scope will be allocated in local memory.

Users can create accessors that reference sub-buffers as well as entire buffers.

Within kernels, the underlying C++ pointer types can be obtained from an accessor. The pointer types will contain a compile-time deduced address space. So, for example, if a C++ pointer is obtained from an accessor to global memory, the C++ pointer type will have a global address space attribute attached to it. The address space attribute will be compile-time propagated to other pointer values when one pointer is initialized to another pointer value using a defined algorithm.

When developers need to explicitly state the address space of a pointer value, one of the explicit pointer classes can be used. There is a different explicit pointer class for each address space: sycl::raw_local_ptr, sycl::raw_global_ptr, sycl::raw_private_ptr, sycl::raw_generic_ptr, sycl::decorated_local_ptr, sycl::decorated_global_ptr, sycl::decorated_private_ptr, or sycl::decorated_generic_ptr.

The classes with the decorated prefix expose pointers that use an implementation-defined address space decoration, while the classes with the raw prefix do not. Buffer accessors with an access target target::device or target::constant_buffer and local accessors can be converted into explicit pointer classes (multi_ptr).

For templates that need to adapt to different address spaces, a sycl::multi_ptr class is defined which is templated via a compile-time constant enumerator value to specify the address space.

3.8.3. SYCL memory consistency model

The SYCL memory consistency model is based upon the memory consistency model of the C++ core language. Where SYCL offers extensions to classes and functions that may affect memory consistency, the default behavior when these extensions are not used always matches the behavior of standard C++.

A SYCL implementation must guarantee that the same memory consistency model is used across host and device code. Every device compiler must support the memory model defined by the minimum version of C++ described in Section 3.9.1; SYCL implementations supporting additional versions of C++ must also support the corresponding memory models.

Within a work-item, operations are ordered according to the sequenced before relation defined by the C++ core language.

Ensuring memory consistency across different work-items requires careful usage of group barrier operations, mem-fence operations and atomic operations. The ordering of operations across different work-items is determined by the happens before relation defined by the C++ core language, with a single relation governing all address spaces (memory regions).

On any SYCL device, local and global memory may be made consistent across work-items in a single group through use of a group barrier operation. On SYCL devices supporting acquire-release or sequentially consistent memory orderings, all memory visible to a set of work-items may be made consistent across the work-items in that set through the use of mem-fence and atomic operations.

Memory consistency between the host and SYCL device(s), or different SYCL devices in the same context, can be guaranteed through library calls in the host application, as defined in Section 3.9.8. On SYCL devices supporting concurrent atomic accesses to USM allocations and acquire-release or sequentially consistent memory orderings, cross-device memory consistency can be enforced through the use of mem-fence and atomic operations.

3.8.3.1. Memory ordering
 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
namespace sycl {

enum class memory_order : /* unspecified */ {
  relaxed,
  acquire,
  release,
  acq_rel,
  seq_cst
};

inline constexpr auto memory_order_relaxed = memory_order::relaxed;
inline constexpr auto memory_order_acquire = memory_order::acquire;
inline constexpr auto memory_order_release = memory_order::release;
inline constexpr auto memory_order_acq_rel = memory_order::acq_rel;
inline constexpr auto memory_order_seq_cst = memory_order::seq_cst;

} // namespace sycl

The memory synchronization order of a given atomic operation is controlled by a sycl::memory_order parameter, which can take one of the following values:

  • sycl::memory_order::relaxed;

  • sycl::memory_order::acquire;

  • sycl::memory_order::release;

  • sycl::memory_order::acq_rel;

  • sycl::memory_order::seq_cst.

The meanings of these values are identical to those defined in the C++ core language.

These memory orders are listed above from weakest (memory_order::relaxed) to strongest (memory_order::seq_cst).

The complete set of memory orders is not guaranteed to be supported by every device, nor across all combinations of devices within a platform. The set of supported memory orders can be queried via the information descriptors for the sycl::device and sycl::context classes.

SYCL implementations are not required to support a memory order equivalent to std::memory_order::consume, and using this ordering within a SYCL device kernel results in undefined behavior. Developers are encouraged to use sycl::memory_order::acquire instead.

3.8.3.2. Memory scope
 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
namespace sycl {

enum class memory_scope : /* unspecified */ {
  work_item,
  sub_group,
  work_group,
  device,
  system
};

inline constexpr auto memory_scope_work_item = memory_scope::work_item;
inline constexpr auto memory_scope_sub_group = memory_scope::sub_group;
inline constexpr auto memory_scope_work_group = memory_scope::work_group;
inline constexpr auto memory_scope_device = memory_scope::device;
inline constexpr auto memory_scope_system = memory_scope::system;

} // namespace sycl

The set of work-items and devices to which the memory ordering constraints of a given atomic operation apply is controlled by a sycl::memory_scope parameter, which can take one of the following values:

  • sycl::memory_scope::work_item The ordering constraint applies only to the calling work-item (this scope is only valid for barriers and fences);

  • sycl::memory_scope::sub_group The ordering constraint applies only to work-items in the same sub-group as the calling work-item;

  • sycl::memory_scope::work_group The ordering constraint applies only to work-items in the same work-group as the calling work-item;

  • sycl::memory_scope::device The ordering constraint applies only to work-items executing on the same device as the calling work-item;

  • sycl::memory_scope::system The ordering constraint applies to any work-item or host thread in the system that is currently permitted to access the memory allocation containing the referenced object, as defined by the capabilities of buffers and USM.

The memory scopes are listed above from narrowest (memory_scope::work_item) to widest (memory_scope::system).

The complete set of memory scopes is not guaranteed to be supported by every device. The set of supported memory scopes can be queried via the information descriptors for the sycl::device and sycl::context classes.

The widest scope that can be applied to an atomic operation corresponds to the set of work-items which can access the associated memory location. For example, the widest scope that can be applied to atomic operations in work-group local memory is sycl::memory_scope::work_group. If a wider scope is supplied, the behavior is as-if the narrowest scope containing all work-items which can access the associated memory location was supplied.

The addition of memory scopes to the C++ memory model modifies the definition of some concepts from the C++ core language. For example: data races, the synchronizes-with relationship and sequential consistency must be defined in a way that accounts for atomic operations with differing (but compatible) scopes, in a manner similar to the OpenCL 2.0 specification. Efforts to formalize the memory model of SYCL are ongoing, and a formal memory model will be included in a future version of the SYCL specification.

3.8.3.3. Atomic operations

Atomic operations can be performed on memory in buffers and USM. The sycl::atomic_ref class must be used to provide safe atomic access to the buffer or USM allocation from device code.

3.8.3.4. Forward progress

This section, and any subsequent section referring to progress guarantees, uses the following terms as defined in the C++ core language: thread of execution; weakly parallel forward progress guarantees; parallel forward progress guarantees; concurrent forward progress guarantees; and block with forward progress guarantee delegation.

Each work-item in SYCL is a separate thread of execution, providing at least weakly parallel forward progress guarantees. Whether work-items provide stronger forward progress guarantees is implementation-defined.

All implementations must additionally ensure that a work-item arriving at a group barrier does not prevent other work-items in the same group from making progress. When a work-item arrives at a group barrier acting on group G, implementations must eventually select and potentially strengthen another work-item in group G that has not yet arrived at the barrier.

When a host thread blocks on the completion of a command previously submitted to a SYCL queue (for example, via the sycl::queue::wait function), it blocks with forward progress guarantee delegation.

SYCL commands submitted to a queue are not guaranteed to begin executing until a host thread blocks on their completion. In the absence of multiple host threads, there is no guarantee that host and device code will execute concurrently.

3.9. The SYCL programming model

A SYCL program is written in standard C++. Host code and device code is written in the same C++ source file, enabling instantiation of templated kernels from host code and also enabling kernel source code to be shared between host and device. The device kernels are encapsulated C++ callable types (a function object with operator() or a lambda expression), which have been designated to be compiled as SYCL kernels.

SYCL programs target heterogeneous systems. The kernels may be compiled and optimized for multiple different processor architectures with very different binary representations.

3.9.1. Minimum version of C++

The C++ features used in SYCL are based on a specific version of C++. Implementations of SYCL must support this minimum C++ version, which defines the C++ constructs that can consequently be used by SYCL feature definitions (for example, lambdas).

The minimum C++ version of this SYCL specification is determined by the normative C++ core language defined in Section 3.3. All implementations of this specification must support at least this core language, and features within this specification are defined using features of the core language. Note that not all core language constructs are supported within SYCL kernel functions or code invoked by a SYCL kernel function, as detailed by Section 5.4.

Implementations may support newer C++ versions than the minimum required by SYCL. Code written using newer features than the SYCL requirement, though, may not be portable to other implementations that don’t support the same C++ version.

3.9.2. Alignment with future versions of C++

Some features of SYCL are aligned with the next C++ specification, as defined in Section 3.3.

The following features are pre-adopted by SYCL 2020 and made available in the sycl:: namespace: std::span, std::dynamic_extent, std::bit_cast. The implementations of pre-adopted features are compliant with the next C++ specification, and are expected to forward directly to standard C++ features in a future version of SYCL.

The following features of SYCL 2020 use syntax based on the next C++ specification: sycl::atomic_ref. These features behave as described in the next C++ specification, barring modifications to ensure compatibility with other SYCL 2020 features and heterogeneous programming. Any such modifications are documented in the corresponding sections of this specification.

3.9.3. Basic data parallel kernels

Data-parallel kernels that execute as multiple work-items and where no work-group-local coordination is required are enqueued with the sycl::parallel_for function parameterized by a sycl::range parameter. These kernels will execute the kernel function body once for each work-item in the specified range.

Functionality tied to groups of work-items, including group barriers and local memory, must not be used within these kernels.

Variables with reduction semantics can be added to basic data parallel kernels using the features described in Section 4.9.2.

3.9.4. Work-group data parallel kernels

Data parallel kernels can also execute in a mode where the set of work-items is divided into work-groups of user-defined dimensions. The user specifies the global range and local work-group size as parameters to the sycl::parallel_for function with a sycl::nd_range parameter. In this mode of execution, kernels execute over the nd-range in work-groups of the specified size. It is possible to share data among work-items within the same work-group in local or global memory, and the group_barrier function can be used to block a work-item until all work-items in the same work-group arrive at the barrier. All work-groups in a given parallel_for will be the same size, and the global size defined in the nd-range must either be a multiple of the work-group size in each dimension, or the global size must be zero. When the global size is zero, the kernel function is not executed, the local size is ignored, and any dependencies are satisfied.

Work-groups may be further subdivided into sub-groups. The work-items that compose a sub-group are selected in an implementation-defined way, and therefore the size and number of sub-groups may differ for each kernel. Moreover, different devices may make different guarantees with respect to how sub-groups within a work-group are scheduled. The maximum number of work-items in any sub-group in a kernel is based on a combination of the kernel and its dispatch dimensions. The size of any sub-group in the dispatch is between 1 and this maximum sub-group size, and the size of an individual sub-group is invariant for the duration of a kernel’s execution. Similarly to work-groups, the group_barrier function can be used to block a work-item until all work-items in the same sub-group arrive at the barrier.

Portable device code must not assume that work-items within a sub-group execute in any particular order, that work-groups are subdivided into sub-groups in a specific way, nor that the work-items within a sub-group provide specific forward progress guarantees.

Variables with reduction semantics can be added to work-group data parallel kernels using the features described in Section 4.9.2.

3.9.5. Hierarchical data parallel kernels (deprecated)

Hierarchical data parallel kernels and all classes that are only available within such kernels are deprecated in SYCL 2020, and will be removed in a future version of SYCL.

The SYCL compiler provides a way of specifying data parallel kernels that execute within work-groups via a different syntax which highlights the hierarchical nature of the parallelism. This mode is purely a compiler feature and does not change the execution model of the kernel. Instead of calling sycl::parallel_for the user calls sycl::parallel_for_work_group with a sycl::range value representing the number of work-groups to launch and optionally a second sycl::range representing the size of each work-group for performance tuning. All code within the parallel_for_work_group scope effectively executes once per work-group. Within the parallel_for_work_group scope, it is possible to call parallel_for_work_item which creates a new scope in which all work-items within the current work-group execute. This enables a programmer to write code that looks like there is an inner work-item loop inside an outer work-group loop, which closely matches the effect of the execution model. All variables declared inside the parallel_for_work_group scope are allocated in work-group local memory, whereas all variables declared inside the parallel_for_work_item scope are declared in private memory. All parallel_for_work_item calls within a given parallel_for_work_group execution must have the same dimensions.

3.9.6. Kernels that are not launched over parallel instances

Simple kernels for which only a single instance of the kernel function will be executed are enqueued with the sycl::single_task function. The kernel enqueued takes no “work-item id” parameter and will only execute once. The behavior is logically equivalent to executing a kernel on a single compute unit with a single work-group comprising only one work-item. Such kernels may be enqueued on multiple queues and devices and as a result may be executed in task-parallel fashion.

3.9.7. Pre-defined kernels

Some SYCL backends may expose pre-defined functionality to users as kernels. These kernels are not programmable, hence they are not bound by the SYCL C++ programming model restrictions, and how they are written is implementation-defined.

3.9.8. Coordination and synchronization

Coordination between the host and any devices can be expressed in the host SYCL application using calls into the SYCL runtime. Coordination between work-items executing inside of device code can be expressed using group barriers.

Some function calls synchronize with other function calls performed by another thread (potentially on another device). Other functions are defined in terms of their synchronization operations. Such functions can be used to ensure that the host and any devices do not access data concurrently, and/or to reason about the ordering of operations across the host and any devices.

3.9.8.1. Host-device coordination

The following operations can be used to coordinate host and device(s):

  • Buffer destruction: The destructors for sycl::buffer, sycl::unsampled_image and sycl::sampled_image objects block until all submitted work on those objects completes and copy the data back to host memory before returning. These destructors only block if the object was constructed with attached host memory and if data needs to be copied back to the host.

    More complex forms of buffer destruction can be specified by the user by constructing buffers with other kinds of references to memory, such as shared_ptr and unique_ptr.

  • Host Accessors: The constructor for a host accessor blocks until all kernels that modify the same buffer (or image) in any queues complete and then copies data back to host memory before the constructor returns. Any command groups with requirements to the same memory object cannot execute until the host accessor is destroyed as shown on Table 6.

  • Command group enqueue: The SYCL runtime internally ensures that any command groups added to queues have the correct event dependencies added to those queues to ensure correct operation. Adding command groups to queues never blocks, and the sycl::event returned by the queue’s submit function contains event information related to the specific command group.

  • Queue operations: The user can manually use queue operations, such as sycl::queue::wait() to block execution of the calling thread until all the command groups submitted to the queue have finished execution. Note that this will also affect the dependencies of those command groups in other queues.

  • SYCL event objects: SYCL provides sycl::event objects which can be used to track and specify dependencies. The SYCL runtime must ensure that these objects can be used to enforce dependencies that span SYCL contexts from different SYCL backends.

The specification for each of these blocking functions defines some set of operations that cause the function to unblock. These operations always happen before the blocking function returns (using the definition of "happens before" from the C++ specification).

Note that the destructors of other SYCL objects (sycl::queue, sycl::context,…) do not block. Only a sycl::buffer, sycl::sampled_image or sycl::unsampled_image destructor might block. The rationale is that an object without any side effect on the host does not need to block on destruction as it would impact the performance. So it is up to the programmer to use a member function to wait for completion in some cases if this does not fit the goal. See Section 3.9.12 for more information on object life time.

3.9.8.2. Work-item coordination

A group barrier provides a mechanism to coordinate all work-items in the same group. All work-items in a group must execute the barrier before any are allowed to continue execution beyond the barrier. Note that the group barrier must be encountered by all work-items of a group executing the kernel or by none at all. work-group barrier and sub-group barrier functionality is exposed via the group_barrier function.

Coordination between work-items in different work-groups must take place via atomic operations, and is possible only on SYCL device with certain capabilities, as described in Section 3.8.3.

3.9.9. Error handling

In SYCL, there are two types of errors: synchronous errors that can be detected immediately when an API call is made, and asynchronous errors that can only be detected later after an API call has returned. Synchronous errors, such as failure to construct an object, are reported immediately by the runtime throwing an exception. Asynchronous errors, such as an error occurring during execution of a kernel on a device, are reported via an asynchronous error-handler mechanism.

Asynchronous errors are not reported immediately as they occur. The asynchronous error handler for a context or queue is called with a sycl::exception_list object, which contains a list of asynchronously-generated exception objects, on the conditions described by Section 4.13.1.1 and Section 4.13.1.2.

Asynchronous errors may be generated regardless of whether the user has specified any asynchronous error handler(s), as described in Section 4.13.1.2.

Some SYCL backends can report errors that are specific to the platform they are targeting, or that are more concrete than the errors provided by the SYCL API. Any error reported by a SYCL backend must derive from the base sycl::exception. When a user wishes to capture specifically an error thrown by a SYCL backend, she must include the SYCL backend-specific headers for said SYCL backend.

3.9.10. Fallback mechanism

A command group function object can be submitted either to a single queue to be executed on, or to a secondary queue. If a command group function object fails to be enqueued to the primary queue, then the implementation may attempt to enqueue it to the secondary queue, if given as a parameter to the submit function. It is implementation defined whether the secondary queue is used as a fallback in this manner. If the command group function object fails to be enqueued to both of these queues, or if it fails to be enqueued to the primary queue and the implementation elects not to enqueue it to the secondary queue, then a synchronous SYCL exception will be thrown.

It is possible that a command group may be successfully enqueued, but then asynchronously fail to run, for some reason. In this case, it may be possible for the runtime system to execute the command group function object on the secondary queue, instead of the primary queue. The situations where a SYCL runtime may be able to achieve this asynchronous fall-back is implementation-defined.

3.9.11. Scheduling of kernels and data movement

A command group function object takes a reference to a command group handler as a parameter and anything within that scope is immediately executed and takes the handler object as a parameter. The intention is that a user will perform calls to SYCL functions, member functions, destructors and constructors inside that scope. These calls will be non-blocking on the host, but enqueue operations to the queue that the command group is submitted to. All user functions within the command group scope will be called on the host as the command group function object is executed, but any commands it invokes will be added to the SYCL queue. All commands added to the queue will be executed out-of-order from each other, according to their data dependencies.

3.9.12. Managing object lifetimes

A SYCL application does not initialize any SYCL backend features until a sycl::context object is created. A user does not need to explicitly create a sycl::context object, but they do need to explicitly create a sycl::queue object, for which a sycl::context object will be implicitly created if not provided by the user.

All SYCL backend objects encapsulated in SYCL objects are reference-counted and will be destroyed once all references have been released. This means that a user needs only create a SYCL queue (which will automatically create an SYCL context) for the lifetime of their application to initialize and release any SYCL backend objects safely.

There is no global state specified to be required in SYCL implementations. This means, for example, that if the user creates two queues without explicitly constructing a common context, then a SYCL implementation does not have to create a shared context for the two queues. Implementations are free to share or cache state globally for performance, but it is not required.

Memory objects can be constructed with or without attached host memory. If no host memory is attached at the point of construction, then destruction of that memory object is non-blocking. The user may use C++ standard pointer classes for sharing the host data with the user application and for defining blocking, or non-blocking behavior of the buffers and images. If host memory is attached by using a raw pointer, then the default behavior is followed, which is that the destructor will block until any command groups operating on the memory object have completed, then, if the contents of the memory object is modified on a device those contents are copied back to host and only then does the destructor return.

In the case where host memory is shared between the user application and the SYCL runtime with a std::shared_ptr, then the reference counter of the std::shared_ptr determines whether the buffer needs to copy data back on destruction, and in that case the blocking or non-blocking behavior depends on the user application.

Instead of a std::shared_ptr, a std::unique_ptr may be provided, which uses move semantics for initializing and using the associated host memory. In this case, the behavior of the buffer in relation to the user application will be non-blocking on destruction.

As said in Section 3.9.8, the only blocking operations in SYCL (apart from explicit wait operations) are:

  • host accessor constructor, which waits for any kernels enqueued before its creation that write to the corresponding object to finish and be copied back to host memory before it starts processing. The host accessor does not necessarily copy back to the same host memory as initially given by the user;

  • memory object destruction, in the case where copies back to host memory have to be done or when the host memory is used as a backing-store.

3.9.13. Device discovery and selection

A user specifies which queue to submit a command group function object and each queue is targeted to run on a specific device (and context). A user can specify the actual device on queue creation, or they can specify a device selector which causes the SYCL runtime to choose a device based on the user’s provided preferences. Specifying a device selector causes the SYCL runtime to perform device discovery. No device discovery is performed until a SYCL device selector is passed to a queue constructor. Device topology may be cached by the SYCL runtime, but this is not required.

Device discovery will return all devices from all platforms exposed by all the supported SYCL backends.

3.9.14. Interfacing with the SYCL backend API

There are two styles of developing a SYCL application:

  1. writing a pure SYCL generic application;

  2. writing a SYCL application that relies on some SYCL backend specific behavior.

When users follow 1., there is no assumption about what SYCL backend will be used during compilation or execution of the SYCL application. Therefore, the SYCL backend API is not assumed to be available to the developer. Only standard C++ types and interfaces are assumed to be available, as described in Section 3.9. Users only need to include the <sycl/sycl.hpp> header to write a SYCL generic application.

On the other hand, when users follow 2., they must know what SYCL backend APIs they are using. In this case, any header required for the normal programmability of the SYCL backend API is assumed to be available to the user. In addition to the <sycl/sycl.hpp> header, users must also include the SYCL backend-specific header as defined in Section 4.3. The SYCL backend-specific header provides the interoperability interface for the SYCL API to interact with native backend objects.

The interoperability API is defined in Section 4.5.1.

3.10. Memory objects

SYCL memory objects represent data that is handled by the SYCL runtime and can represent allocations in one or multiple devices at any time. Memory objects, both buffers and images, may have one or more underlying native backend objects to ensure that queues objects can use data in any device. A SYCL implementation may have multiple native backend objects for the same device. The SYCL runtime is responsible for ensuring the different copies are up-to-date whenever necessary, using whatever mechanism is available in the system to update the copies of the underlying native backend objects.

Implementation note

A valid mechanism for this update is to transfer the data from one SYCL backend into the system memory using the SYCL backend-specific mechanism available, and then transfer it to a different device using the mechanism exposed by the new SYCL backend.

Memory objects in SYCL fall into one of two categories: buffer objects and image objects. A buffer object stores a one-, two- or three-dimensional collection of elements that are stored linearly directly back to back in the same way C or C++ stores arrays. An image object is used to store a one-, two- or three-dimensional texture, frame-buffer or image data that may be stored in an optimized and device-specific format in memory and must be accessed through specialized operations.

Elements of a buffer object can be a scalar data type (such as an int or float), vector data type, or a user-defined structure. In SYCL, a buffer object is a templated type (sycl::buffer), parameterized by the element type and number of dimensions. An image object is stored in one of a limited number of formats. The elements of an image object are selected from a list of predefined image formats which are provided by an underlying SYCL backend implementation. Images are encapsulated in the sycl::unsampled_image or sycl::sampled_image types, which are templated by the number of dimensions in the image. The minimum number of elements in an image object is one. The minimum number of elements in a buffer object is zero.

The fundamental differences between a buffer and an image object are:

  • elements in a buffer are stored in an array of 1, 2 or 3 dimensions and can be accessed using an accessor by a kernel executing on a device. The accessors for kernels provide a member function to get C++ pointer types, or the sycl::global_ptr class;

  • elements of an image are stored in a format that is opaque to the user and cannot be directly accessed using a pointer. SYCL provides image accessors and samplers to allow a kernel to read from or write to an image;

  • for a buffer object the data is accessed within a kernel in the same format as it is stored in memory, but in the case of an image object the data is not necessarily accessed within a kernel in the same format as it is stored in memory;

  • image elements are always a 4-component vector (each component can be a float or signed/unsigned integer) in a kernel. Accessors that read an image convert image elements from their storage format into a 4-component vector.

    Similarly, the SYCL accessor member functions provided to write to an image convert the image element from a 4-component vector to the appropriate image format specified such as four 8-bit elements, for example.

Users may want fine-grained control of the memory management and storage semantics of SYCL image or buffer objects. For example, a user may wish to specify the host memory for a memory object to use, but may not want the memory object to block on destruction.

Depending on the control and the use cases of the SYCL applications, well established C++ classes and patterns can be used for reference counting and sharing data between user applications and the SYCL runtime. For control over memory allocation on the host and mapping between host and device memory, pre-defined or user-defined C++ std::allocator classes are used. To avoid data races when sharing data between SYCL and non-SYCL applications, std::shared_ptr and std::mutex classes are used.

3.11. Multi-dimensional objects and linearization

SYCL defines a number of multi-dimensional objects such as buffers and accessors. The iteration space of work-items in a kernel may also be multi-dimensional. The size of each dimension is defined by a range object of one, two or three dimensions, and an element in the multi-dimensional space can be identified using an id object with the same number of dimensions as the corresponding range.

If the size of any dimension is zero, there are zero elements in the multi-dimensional range.

3.11.1. Linearization

Some multi-dimensional objects can be viewed in a linear form. When this happens, the right-most term in the object’s range varies fastest in the linearization.

A three-dimensional element id{id0, id1, id2} within a three-dimensional object of range range{r0, r1, r2} has a linear position defined by:

A two-dimensional element id{id0, id1} within a two-dimensional range{r0, r1} follows a similar equation:

A one-dimensional element id{id0} within a one-dimensional range range{r0} is equivalent to its linear form.

3.11.2. Multi-dimensional subscript operators

Some multi-dimensional objects can be indexed using the subscript operator where consecutive subscript operators correspond to each dimension. The right-most operator varies fastest, as with standard C++ arrays. Formally, a three-dimensional subscript access a[id0][id1][id2] references the element at id{id0, id1, id2}. A two-dimensional subscript access a[id0][id1] references the element at id{id0, id1}. A one-dimensional subscript access a[id0] references the element at id{id0}.

3.12. Implementation options

The SYCL language is designed to allow several different possible implementations. The contents of this section are non-normative, so implementations need not follow the guidelines listed here. However, this section is intended to help readers understand the possible strategies that can be used to implement SYCL.

3.12.1. Single source multiple compiler passes

With this technique, known as SMCP, there are separate host and device compilers. Each SYCL source file is compiled two times: once by the host compiler and once by the device compiler. An implementation could support more than one device compiler, in which case each SYCL source file is compiled more than two times. The host compiler in this technique could be an off-the-shelf compiler with no special knowledge of SYCL, but the device compiler must be SYCL aware. The device compiler parses the source file to identify each SYCL kernel function and any device functions it calls. SYCL is designed so that this analysis can be done statically. The device compiler then generates code only for the SYCL kernel functions and the device functions.

Typically, the device compilers generate header files which interface between the host compiler and the SYCL runtime. Therefore, the device compiler runs first, and then the host compiler consumes these header files when generating the host code.

The device compilers in this technique generate one or more device images for the SYCL kernel functions, which can be read by the SYCL runtime. Each device image could either contain native ISA for a device or it could contain an intermediate language such as SPIR-V. In the later case, the SYCL runtime must translate the intermediate language into native device ISA when the SYCL kernel function is submitted to a device.

Since this technique has separate host and device compilers, there needs to be some way to associate a SYCL kernel function (which is compiled by the device compiler) with the code that invokes it (which is compiled by the host compiler). Implementations conformant to the reduced feature set (Section B.2) can do this by using the C++ type of the SYCL kernel function. This type is specified via the kernel name template parameter if the SYCL kernel function is a lambda expression, or it is obtained from the class type if the SYCL kernel function is an object. Implementations conformant to the full feature set (Section B.1) do not require a kernel name at the invocation site, so they must implement some other way to make the association.

3.12.2. Single source single compiler pass

With this technique, known as SSCP, the vendor implements a custom compiler that reads each SYCL source file only once, and that compiler generates the host code as well as the device images for the SYCL kernel functions. As in the SMCP case, each device image could either contain native device ISA or an intermediate language.

3.12.3. Library-only implementation

It is also possible to implement SYCL purely as a library, using an off-the-shelf host compiler with no special support for SYCL. In such an implementation, each kernel may run on the host system.

3.13. Language restrictions in kernels

The SYCL kernels are executed on SYCL devices and all of the functions called from a SYCL kernel are going to be compiled for the device by a SYCL device compiler. Due to restrictions of the heterogeneous devices where the SYCL kernel will execute, there are certain restrictions on the base C++ language features that can be used inside kernel code. For details on language restrictions please refer to Section 5.4.

SYCL kernels use arguments that are captured by value in the command group scope or are passed from the host to the device using accessors. Sharing data structures between host and device code imposes certain restrictions, such as using only objects that are device copyable, and in general, no pointers initialized for the host can be used on the device. SYCL memory objects, such as sycl::buffer, sycl::unsampled_image, and sycl::sampled_image, cannot be passed to a kernel. Instead, a kernel must interact with these objects through accessors. No hierarchical structures of these memory object classes are supported and any other data containers need to be converted to the SYCL data management classes using the SYCL interface. For more details on the rules for kernel parameter passing, please refer to Section 4.12.4.

Pointers to USM allocations may be passed to a kernel either directly as arguments or indirectly inside of other objects. Pointers to USM allocations that are passed as kernel arguments are treated as being in the global address space.

3.13.1. Device copyable

The SYCL implementation may need to copy data between the host and a device or between two devices. For example, this may occur when a command group has a requirement for the contents of a buffer or when the application passes certain arguments to a SYCL kernel function (as described in Section 4.12.4). Such data must have a type that is device copyable, as defined below.

An implementation can assume that it is always safe to perform bitwise copies of any object that has a device copyable type.

Any type that is trivially copyable (as defined by the C++ core language) is implicitly device copyable.

Although implementations are not required to support device code that calls library functions from the C++ core language, some implementations may provide device support for some of these functions. If the implementation provides device support for one of the following classes, that type is also implicitly device copyable:

  • std::array<T, 0>;

  • std::array<T, N> if T is device copyable;

  • std::optional<T> if T is device copyable;

  • std::pair<T1, T2> if T1 and T2 are device copyable;

  • std::tuple<>;

  • std::tuple<Types...> if all the types in the parameter pack Types are device copyable;

  • std::variant<>;

  • std::variant<Types...> if all the types in the parameter pack Types are device copyable;

  • std::basic_string_view<CharT, Traits>;

  • std::span<ElementType, Extent> (the std::span type has been introduced in C++20);

  • sycl::span<ElementType, Extent>.

If the implementation provides device support for one of the classes listed above, arrays of that class and cv-qualified versions of that class are also device copyable.

Types such as std::basic_string_view<CharT, Traits> and std::span<ElementType, Extent> are view types, which reference underlying data that they do not own. Copying such a type only copies the view and not the referenced data. If a view is copied between the host and device or between two devices, it is the application’s responsibility to ensure that the referenced data is allocated in memory that can be accessed by the recipient (see Section 4.8).

In addition, the implementation may allow the application to explicitly declare certain class types as device copyable. If the implementation has this support, it must predefine the preprocessor macro SYCL_DEVICE_COPYABLE to 1, and it must not predefine this preprocessor macro if it does not have this support. When the implementation has this support, an application may declare that a class type T is device copyable by defining the trait is_device_copyable_v<T> to true if all of the following statements are true:

  • Type T has at least one eligible copy constructor, move constructor, copy assignment operator, or move assignment operator;

  • Each eligible copy constructor, move constructor, copy assignment operator, and move assignment operator is public;

  • The effect of each eligible copy constructor, move constructor, copy assignment operator, and move assignment operator is the same as a bitwise copy of the object;

  • Type T has a public non-deleted destructor; and

  • The destructor has no effect.

Declaring that a class type T is device copyable when any of these statements is not true results in undefined behavior.

When the application explicitly declares a class type to be device copyable, arrays of that type and cv-qualified versions of that type are also device copyable, and the implementation sets the is_device_copyable_v trait to true for these array and cv-qualified types.

3.14. Endianness support

SYCL does not mandate any particular byte order, but the byte order of the host always matches the byte order of the devices. This allows data to be copied between the host and the devices without any byte swapping.

3.15. Example SYCL application

Below is a more complex example application, combining some of the features described above.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
#include <iostream>
#include <sycl/sycl.hpp>
using namespace sycl;  // (optional) avoids need for "sycl::" before SYCL names

// Size of the matrices
constexpr size_t N = 2000;
constexpr size_t M = 3000;

int main() {
  // Create a queue to work on
  queue myQueue;

  // Create some 2D buffers of float for our matrices
  buffer<float, 2> a{range<2>{N, M}};
  buffer<float, 2> b{range<2>{N, M}};
  buffer<float, 2> c{range<2>{N, M}};

  // Launch an asynchronous kernel to initialize a
  myQueue.submit([&](handler& cgh) {
    // The kernel writes a, so get a write accessor on it
    accessor A{a, cgh, write_only};

    // Enqueue a parallel kernel iterating on a N*M 2D iteration space
    cgh.parallel_for(range<2>{N, M},
                     [=](id<2> index) { A[index] = index[0] * 2 + index[1]; });
  });

  // Launch an asynchronous kernel to initialize b
  myQueue.submit([&](handler& cgh) {
    // The kernel writes b, so get a write accessor on it
    accessor B{b, cgh, write_only};

    // From the access pattern above, the SYCL runtime detects that this
    // command_group is independent from the first one and can be
    // scheduled independently

    // Enqueue a parallel kernel iterating on a N*M 2D iteration space
    cgh.parallel_for(range<2>{N, M}, [=](id<2> index) {
      B[index] = index[0] * 2014 + index[1] * 42;
    });
  });

  // Launch an asynchronous kernel to compute matrix addition c = a + b
  myQueue.submit([&](handler& cgh) {
    // In the kernel a and b are read, but c is written
    accessor A{a, cgh, read_only};
    accessor B{b, cgh, read_only};
    accessor C{c, cgh, write_only};

    // From these accessors, the SYCL runtime will ensure that when
    // this kernel is run, the kernels computing a and b have completed

    // Enqueue a parallel kernel iterating on a N*M 2D iteration space
    cgh.parallel_for(range<2>{N, M},
                     [=](id<2> index) { C[index] = A[index] + B[index]; });
  });

  // Ask for an accessor to read c from application scope.  The SYCL runtime
  // waits for c to be ready before returning from the constructor
  host_accessor C{c, read_only};
  std::cout << std::endl << "Result:" << std::endl;
  for (size_t i = 0; i < N; i++) {
    for (size_t j = 0; j < M; j++) {
      // Compare the result to the analytic value
      if (C[i][j] != i * (2 + 2014) + j * (1 + 42)) {
        std::cout << "Wrong value " << C[i][j] << " on element " << i << " "
                  << j << std::endl;
        exit(-1);
      }
    }
  }

  std::cout << "Good computation!" << std::endl;
  return 0;
}

4. SYCL programming interface

The SYCL programming interface provides a common abstracted feature set to one or more SYCL backend APIs. This section describes the C++ library interface to the SYCL runtime which executes across those SYCL backends.

The entirety of the SYCL interface defined in this section is required to be available for any SYCL backends, with the exception of the interoperability interface, which is described in general terms in this document, not pertaining to any particular SYCL backend.

All functions defined in this specification are thread-safe, unless otherwise specified.

The underlying types for all enumerations defined in this specification are implementation-defined. In addition, all enumerators within an enumeration have some implementation-defined unique value unless the specification specifically indicates a value for the enumerator.

4.1. Backends

The SYCL backends that can be supported by a SYCL implementation are identified using the enum class backend.

1
2
3
4
5
namespace sycl {
enum class backend : /* unspecified */ {
  /* see below */
};
} // namespace sycl

The enum class backend is implementation-defined and must be populated with a unique identifier for each SYCL backend that the SYCL implementation can support. Note that the SYCL backends listed in the enum class backend are not guaranteed to be available in a given installation.

Each named SYCL backend enumerated in the enum class backend must be associated with a SYCL backend specification. Many sections of this specification will refer to the associated SYCL backend specification.

4.1.1. Backend macros

As the identifiers defined in enum class backend are implementation-defined, and the associated backends are not guaranteed to be available, a SYCL implementation must also define a preprocessor macro for each of these identifiers. If the SYCL backend is defined by the Khronos SYCL group, the name of the macro has the form SYCL_BACKEND_<backend_name>, where backend_name is the associated identifier from backend in all upper-case. See Chapter 6 for the name of the macro if the vendor defines the SYCL backend outside of the Khronos SYCL group.

If a backend listed in the enum class backend is not available, the associated macro must be left undefined.

4.2. Generic vs non-generic SYCL

The SYCL programming API is split into two categories; generic SYCL and non-generic SYCL. Almost everything in the SYCL programming API is considered generic SYCL. However any usage of the enum class backend is considered non-generic SYCL and should only be used for SYCL backend specialized code paths, as the identifiers defined in backend are implementation-defined.

In any non-generic SYCL application code where the backend enum class is used, the expression must be guarded with a preprocessor #ifdef guard using the associated preprocessor macro to ensure that the SYCL application will compile even if the SYCL implementation does not support that SYCL backend being specialized for.

4.3. Header files and namespaces

SYCL provides one standard header file: <sycl/sycl.hpp>, which needs to be included in every translation unit that uses the SYCL programming API.

All SYCL classes, constants, types and functions defined by this specification should exist within the ::sycl namespace.

For compatibility with SYCL 1.2.1, SYCL provides another standard header file: <CL/sycl.hpp>, which can be included in place of <sycl/sycl.hpp>. In that case, all SYCL classes, constants, types and functions defined by this specification should exist within the ::cl::sycl C++ namespace. The <CL/sycl.hpp> header and all declarations within the ::cl::sycl namespace are deprecated.

For consistency, the programming API will only refer to the <sycl/sycl.hpp> header and the ::sycl namespace, but this should be considered synonymous with the SYCL 1.2.1 header and namespace.

Include paths starting with "sycl/ext/" and "sycl/backend/" are reserved for extensions to SYCL and for backend interop headers respectively. Other include paths starting with "sycl/" and the sycl::detail namespace are reserved for implementation details.

When a SYCL backend is defined by the Khronos SYCL group, functionality for that SYCL backend is available via the header "sycl/backend/<backend_name>.hpp", and all SYCL backend-specific functionality is made available in the namespace sycl::<backend_name> where <backend_name> is the name of the SYCL backend as defined in the SYCL backend specification.

Chapter 6 defines the allowable header files and namespaces for any extensions that a vendor may provide, including any SYCL backend that the vendor may define outside of the Khronos SYCL group.

Unless otherwise specified, the behavior of a SYCL program is undefined if:

  • it adds any entity to namespace sycl or to a namespace within namespace sycl; or

  • it adds a template specialization for a class template defined in namespace sycl or defined in a namespace within namespace sycl.

4.4. Class availability

In SYCL some SYCL runtime classes are available to the SYCL application, some are available within a SYCL kernel function and some are available on both and can be passed as arguments to a SYCL kernel function.

Each of the following SYCL runtime classes: buffer, buffer_allocator, context, device, device_image, event, exception, handler, host_accessor, host_sampled_image_accessor, host_unsampled_image_accessor, id, image_allocator, kernel, kernel_id, marray, kernel_bundle, nd_range, platform, queue, range, sampled_image, image_sampler, stream, unsampled_image and vec must be available to the host application.

Each of the following SYCL runtime classes: accessor, atomic_ref, device_event, group, h_item, id, item, local_accessor, marray, multi_ptr, nd_item, range, reducer, sampled_image_accessor, stream, sub_group, unsampled_image_accessor and vec must be available within a SYCL kernel function.

4.5. Common interface

When a dimension template parameter is used in SYCL classes, it is defaulted as 1 in most cases.

4.5.1. Backend interoperability

Many of the SYCL runtime classes may be implemented such that they encapsulate an object unique to the SYCL backend that underpins the functionality of that class. Where appropriate, these classes may provide an interface for interoperating between the SYCL runtime object and the native backend object in order to support interoperability within an application between SYCL and the associated SYCL backend API.

There are three forms of interoperability with SYCL runtime classes: interoperability on the SYCL application with the SYCL backend API, interoperability within a SYCL kernel function with the equivalent kernel language types of the SYCL backend, and interoperability within a host task with the interop_handle.

SYCL application interoperability, SYCL kernel function interoperability and host task interoperability are provided via different interfaces and may have different behavior for the same SYCL object.

SYCL application interoperability may be provided for buffer, context, device, device_image, event, kernel, kernel_bundle, platform, queue, sampled_image, and unsampled_image.

SYCL kernel function interoperability may be provided for accessor, device_event, local_accessor, sampled_image_accessor, stream and unsampled_image_accessor inside kernel scope only and is not available outside of that scope.

host task interoperability may be provided for accessor, sampled_image_accessor, unsampled_image_accessor, queue, device, context inside the scope of a host task only, see Section 4.10.

Support for SYCL backend interoperability is optional and therefore not required to be provided by a SYCL implementation. A SYCL application using SYCL backend interoperability is considered to be non-generic SYCL.

Details on the interoperability for a given SYCL backend are available on the SYCL backend specification document for that SYCL backend.

4.5.1.1. Type traits backend_traits
 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
namespace sycl {

template <backend Backend> class backend_traits {
 public:
  template <class T> using input_type = /* see below */;

  template <class T> using return_type = /* see below */;
};

template <backend Backend, typename SyclType>
using backend_input_t =
    typename backend_traits<Backend>::template input_type<SyclType>;

template <backend Backend, typename SyclType>
using backend_return_t =
    typename backend_traits<Backend>::template return_type<SyclType>;

} // namespace sycl

A series of type traits are provided for SYCL backend interoperability, defined in the backend_traits class.

A specialization of backend_traits must be provided for each named SYCL backend enumerated in the enum class backend that is available at compile time.

The type alias backend_input_t is provided to enable less verbose access to the input_type type within backend_traits for a specific SYCL object of type T. The type alias backend_return_t is provided to enable less verbose access to the return_type type within backend_traits for a specific SYCL object of type T.

4.5.1.2. Template function get_native
1
2
3
4
5
6
namespace sycl {

template <backend Backend, class T>
backend_return_t<Backend, T> get_native(const T& syclObject);

} // namespace sycl

For each SYCL runtime class T which supports SYCL application interoperability, a specialization of get_native must be defined, which takes an instance of T and returns a SYCL application interoperability native backend object associated with syclObject which can be used for SYCL application interoperability. The lifetime of the object returned is backend-defined and specified in the backend specification.

For each SYCL runtime class T which supports kernel function interoperability, a specialization of get_native must be defined, which takes an instance of T and returns the kernel function interoperability native backend object associated with syclObject which can be used for kernel function interoperability. The availability and behavior of these template functions are defined by the SYCL backend specification document.

In host code, the get_native function must throw a synchronous exception with the errc::backend_mismatch error code if the backend of the SYCL object doesn’t match the target backend. In device code, the behavior is undefined if the backend of the SYCL object doesn’t match the target backend.

4.5.1.3. Template functions make_*
 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
namespace sycl {

template <backend Backend>
platform make_platform(const backend_input_t<Backend, platform>& backendObject);

template <backend Backend>
device make_device(const backend_input_t<Backend, device>& backendObject);

template <backend Backend>
context make_context(const backend_input_t<Backend, context>& backendObject,
                     const async_handler asyncHandler = {});

template <backend Backend>
queue make_queue(const backend_input_t<Backend, queue>& backendObject,
                 const context& targetContext,
                 const async_handler asyncHandler = {});

template <backend Backend>
event make_event(const backend_input_t<Backend, event>& backendObject,
                 const context& targetContext);

template <backend Backend, typename T, int Dimensions = 1,
          typename AllocatorT = buffer_allocator<std::remove_const_t<T>>>
buffer<T, Dimensions, AllocatorT>
make_buffer(const backend_input_t<Backend, buffer<T, Dimensions, AllocatorT>>&
                backendObject,
            const context& targetContext, event availableEvent);

template <backend Backend, typename T, int Dimensions = 1,
          typename AllocatorT = buffer_allocator<std::remove_const_t<T>>>
buffer<T, Dimensions, AllocatorT>
make_buffer(const backend_input_t<Backend, buffer<T, Dimensions, AllocatorT>>&
                backendObject,
            const context& targetContext);

template <backend Backend, int Dimensions = 1,
          typename AllocatorT = sycl::image_allocator>
sampled_image<Dimensions, AllocatorT> make_sampled_image(
    const backend_input_t<Backend, sampled_image<Dimensions, AllocatorT>>&
        backendObject,
    const context& targetContext, image_sampler imageSampler,
    event availableEvent);

template <backend Backend, int Dimensions = 1,
          typename AllocatorT = sycl::image_allocator>
sampled_image<Dimensions, AllocatorT> make_sampled_image(
    const backend_input_t<Backend, sampled_image<Dimensions, AllocatorT>>&
        backendObject,
    const context& targetContext, image_sampler imageSampler);

template <backend Backend, int Dimensions = 1,
          typename AllocatorT = sycl::image_allocator>
unsampled_image<Dimensions, AllocatorT> make_unsampled_image(
    const backend_input_t<Backend, unsampled_image<Dimensions, AllocatorT>>&
        backendObject,
    const context& targetContext, event availableEvent);

template <backend Backend, int Dimensions = 1,
          typename AllocatorT = sycl::image_allocator>
unsampled_image<Dimensions, AllocatorT> make_unsampled_image(
    const backend_input_t<Backend, unsampled_image<Dimensions, AllocatorT>>&
        backendObject,
    const context& targetContext);

template <backend Backend, bundle_state State>
kernel_bundle<State> make_kernel_bundle(
    const backend_input_t<Backend, kernel_bundle<State>>& backendObject,
    const context& targetContext);

template <backend Backend>
kernel make_kernel(const backend_input_t<Backend, kernel>& backendObject,
                   const context& targetContext);

} // namespace sycl

For each SYCL runtime class T which supports SYCL application interoperability, a specialization of the appropriate template function make_{sycl_class} where {sycl_class} is the class name of T, must be defined, which takes a SYCL application interoperability native backend object and constructs and returns an instance of T. The availability and behavior of these template functions are defined by the SYCL backend specification document.

Overloads of the make_{sycl_class} function which take a SYCL context object as an argument must throw an exception with the errc::backend_mismatch error code if the backend of the provided SYCL context doesn’t match the target backend.

4.5.2. Common reference semantics

Each of the following SYCL runtime classes: accessor, buffer, context, device, device_image, event, host_accessor, host_sampled_image_accessor, host_unsampled_image_accessor, kernel, kernel_id, kernel_bundle, local_accessor, platform, queue, sampled_image, sampled_image_accessor, stream, unsampled_image and unsampled_image_accessor must obey the following statements, where T is the runtime class type:

  • T must be copy constructible and copy assignable in the host application and within SYCL kernel functions in the case that T is a valid kernel argument. Any instance of T that is constructed as a copy of another instance, via either the copy constructor or copy assignment operator, must behave as-if it were the original instance and as-if any action performed on it were also performed on the original instance and must represent the same underlying native backend object as the original instance where applicable.

  • T must be destructible in the host application and within SYCL kernel functions in the case that T is a valid kernel argument. When any instance of T is destroyed, including as a result of the copy assignment operator, any behavior specific to T that is specified as performed on destruction is only performed if this instance is the last remaining host copy, in accordance with the above definition of a copy.

  • T must be move constructible and move assignable in the host application and within SYCL kernel functions in the case that T is a valid kernel argument. Any instance of T that is constructed as a move of another instance, via either the move constructor or move assignment operator, must replace the original instance rendering said instance invalid and must represent the same underlying native backend object as the original instance where applicable.

  • T must be equality comparable in the host application. Equality between two instances of T (i.e. a == b) must be true if one instance is a copy of the other and non-equality between two instances of T (i.e. a != b) must be true if neither instance is a copy of the other, in accordance with the above definition of a copy, unless either instance has become invalidated by a move operation. By extension of the requirements above, equality on T must guarantee to be reflexive (i.e. a == a), symmetric (i.e. a == b implies b == a and a != b implies b != a) and transitive (i.e. a == b && b == c implies c == a).

  • A specialization of std::hash for T must exist in the host application that returns a unique value such that if two instances of T are equal, in accordance with the above definition, then their resulting hash values are also equal and subsequently if two hash values are not equal, then their corresponding instances are also not equal, in accordance with the above definition.

Some SYCL runtime classes will have additional behavior associated with copy, movement, assignment or destruction semantics. If these are specified they are in addition to those specified above unless stated otherwise.

Each of the runtime classes mentioned above must provide a common interface of special member functions in order to fulfill the copy, move, destruction requirements and hidden friend functions in order to fulfill the equality requirements.

A hidden friend function is a function first declared via a friend declaration with no additional out of class or namespace scope declarations. Hidden friend functions are only visible to ADL (Argument Dependent Lookup) and are hidden from qualified and unqualified lookup. Hidden friend functions have the benefits of avoiding accidental implicit conversions and faster compilation.

These common special member functions and hidden friend functions are described in Table 7 and Table 8 respectively.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
namespace sycl {

class T {
  ...

      public : T(const T& rhs);

  T(T&& rhs);

  T& operator=(const T& rhs);

  T& operator=(T&& rhs);

  ~T();

  ...

      friend bool
      operator==(const T& lhs, const T& rhs) { /* ... */
  }

  friend bool operator!=(const T& lhs, const T& rhs) { /* ... */ }

  ...
};
} // namespace sycl
Table 7. Common special member functions for reference semantics
Special member function Description
T(const T& rhs)

Constructs a T instance as a copy of the RHS SYCL T in accordance with the requirements set out above.

T(T&& rhs)

Constructs a SYCL T instance as a move of the RHS SYCL T in accordance with the requirements set out above.

T& operator=(const T& rhs)

Assigns this SYCL T instance with a copy of the RHS SYCL T in accordance with the requirements set out above.

T& operator=(T&& rhs)

Assigns this SYCL T instance with a move of the RHS SYCL T in accordance with the requirements set out above.

~T()

Destroys this SYCL T instance in accordance with the requirements set out in Section 4.5.2. On destruction of the last copy, may perform additional lifetime related operations required for the underlying native backend object specified in the SYCL backend specification document, if this SYCL T instance was originally constructed using one of the backend interoperability make_* functions specified in Section 4.5.1.3. See the relevant backend specification for details.

Table 8. Common hidden friend functions for reference semantics
Hidden friend function Description
bool operator==(const T& lhs, const T& rhs)

Returns true if this LHS SYCL T is equal to the RHS SYCL T in accordance with the requirements set out above, otherwise returns false.

bool operator!=(const T& lhs, const T& rhs)

Returns true if this LHS SYCL T is not equal to the RHS SYCL T in accordance with the requirements set out above, otherwise returns false.

4.5.3. Common by-value semantics

Each of the following SYCL runtime classes: id, range, item, nd_item, h_item, group, sub_group and nd_range must follow the following statements, where T is the runtime class type:

  • T must be copy constructible and copy assignable in the host application (in the case where T is available on the host) and within SYCL kernel functions.

  • T must be destructible in the host application (in the case where T is available on the host) and within SYCL kernel functions.

  • T must be move constructible and move assignable in the host application (in the case where T is available on the host) and within SYCL kernel functions.

  • T must be equality comparable in the host application (in the case where T is available on the host) and within SYCL kernel functions. Equality between two instances of T (i.e. a == b) must be true if the value of all members are equal and non-equality between two instances of T (i.e. a != b) must be true if the value of any members are not equal, unless either instance has become invalidated by a move operation. By extension of the requirements above, equality on T must guarantee to be reflexive (i.e. a == a), symmetric (i.e. a == b implies b == a and a != b implies b != a) and transitive (i.e. a == b && b == c implies c == a).

Some SYCL runtime classes will have additional behavior associated with copy, movement, assignment or destruction semantics. If these are specified they are in addition to those specified above unless stated otherwise.

Each of the runtime classes mentioned above must provide a common interface of special member functions and member functions in order to fulfill the copy, move, destruction and equality requirements, following the rule of five and the rule of zero.

These common special member functions and hidden friend functions are described in Table 9 and Table 10 respectively.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
namespace sycl {

class T {
  ...

      public
      :
      // If any of the following five special member functions are declared,
      // then all five of them should be explicitly declared (see rule of
      // five).
      //
      // Otherwise, none of them should be explicitly declared
      // (see rule of zero).

      // T(const T &rhs);

      // T(T &&rhs);

      // T &operator=(const T &rhs);

      // T &operator=(T &&rhs);

      // ~T();

      ...

      friend bool
      operator==(const T& lhs, const T& rhs) { /* ... */
  }

  friend bool operator!=(const T& lhs, const T& rhs) { /* ... */ }

  ...
};
} // namespace sycl
Table 9. Common special member functions for by-value semantics
Special member function (see rule of five and rule of zero) Description
T(const T& rhs);

Copy constructor.

T(T&& rhs);

Move constructor.

T& operator=(const T& rhs);

Copy assignment operator.

T& operator=(T&& rhs);

Move assignment operator.

~T();

Destructor.

Table 10. Common hidden friend functions for by-value semantics
Hidden friend function Description
bool operator==(const T& lhs, const T& rhs)

Returns true if this LHS SYCL T is equal to the RHS SYCL T in accordance with the requirements set out above, otherwise returns false.

bool operator!=(const T& lhs, const T& rhs)

Returns true if this LHS SYCL T is not equal to the RHS SYCL T in accordance with the requirements set out above, otherwise returns false.

4.5.4. Properties

Each of the following SYCL runtime classes: accessor, buffer, host_accessor, host_sampled_image_accessor, host_unsampled_image_accessor, context, local_accessor, queue, sampled_image, sampled_image_accessor, stream, unsampled_image, unsampled_image_accessor and usm_allocator provide an optional parameter in each of their constructors to provide a property_list which contains zero or more properties. Each of those properties augments the semantics of the class with a particular feature. Each of those classes must also provide has_property and get_property member functions for querying for a particular property.

The listing below illustrates the usage of various buffer properties, described in Section 4.7.2.3.

The example illustrates how using properties does not affect the type of the object, thus, does not prevent the usage of SYCL objects in containers.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
{
  context myContext;

  std::vector<buffer<int, 1>> bufferList{
      buffer<int, 1>{ptr, rng},
      buffer<int, 1>{ptr, rng, property::use_host_ptr{}},
      buffer<int, 1>{ptr, rng, property::context_bound{myContext}}};

  for (auto& buf : bufferList) {
    if (buf.has_property<property::context_bound>()) {
      auto prop = buf.get_property<property::context_bound>();
      assert(myContext == prop.get_context());
    }
  }
}

Each property is represented by a unique class and an instance of a property is an instance of that type. Some properties can be default constructed while others will require an argument on construction. A property may be applicable to more than one class, however some properties may not be compatible with each other. See the requirements for the properties of the SYCL buffer class, SYCL unsampled_image class and SYCL sampled_image class in Section 4.7.2.3 and Table 20 respectively.

Properties can be passed to a SYCL runtime class via an instance of property_list. These properties get tied to the SYCL runtime class instance and copies of the object will contain the same properties.

A SYCL implementation or a SYCL backend may provide additional properties other than those defined here, provided they are defined in accordance with the requirements described in Section 4.3.

4.5.4.1. Properties interface

Each of the runtime classes mentioned above must provide a common interface of member functions in order to fulfill the property interface requirements.

A synopsis of the common properties interface, the SYCL property_list class and the SYCL property classes is provided below. The member functions of the common properties interface are listed in Table 12. The constructors of the SYCL property_list class are listed in Table 13.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
namespace sycl {

template <typename Property> struct is_property;

template <typename Property>
inline constexpr bool is_property_v = is_property<Property>::value;

template <typename Property, typename SyclObject> struct is_property_of;

template <typename Property, typename SyclObject>
inline constexpr bool is_property_of_v =
    is_property_of<Property, SyclObject>::value;

class T {
  ...

  template <typename Property>
  bool has_property() const noexcept;

  template <typename Property> Property get_property() const;

  ...
};

class property_list {
 public:
  template <typename... Properties> property_list(Properties... props);
};
} // namespace sycl
Table 11. Traits for properties
Traits Description
template <typename Property> struct is_property

An explicit specialization of is_property that inherits from std::true_type must be provided for each property, where Property is the class defining the property. This includes both standard properties described in this specification and any additional non-standard properties defined by an implementation. All other specializations of is_property must inherit from std::false_type.

template <typename Property>
inline constexpr bool is_property_v;

Variable containing value of is_property<Property>.

template <typename Property, SyclObject> struct is_property_of

An explicit specialization of is_property_of that inherits from std::true_type must be provided for each property that can be used in constructing a given SYCL class, where Property is the class defining the property and SyclObject is the SYCL class. This includes both standard properties described in this specification and any additional non-standard properties defined by an implementation. All other specializations of is_property_of must inherit from std::false_type.

template <typename Property, SyclObject>
inline constexpr bool is_property_of_v;

Variable containing value of is_property_of<Property, SyclObject>.

Table 12. Common member functions of the SYCL property interface
Member function Description
template <typename Property> bool has_property() const noexcept

Returns true if T was constructed with the property specified by Property. Returns false if it was not.

template <typename Property> Property get_property() const

Returns a copy of the property of type Property that T was constructed with. Must throw an exception with the errc::invalid error code if T was not constructed with the Property property.

Table 13. Constructors of the SYCL property_list class
Constructor Description
template <typename... PropertyN> property_list(PropertyN... props)

Available only when: is_property<property>::value evaluates to true where property is each property in PropertyN.

Construct a SYCL property_list with zero or more properties.

4.5.5. Information queries

Several classes in SYCL provide a generic mechanism for querying the class for information.

Each available query is described by an information descriptor, which is a class or class template that encapsulates an information query and its return type.

4.5.5.1. Information query interface

The information query interface consists of two function templates, templated on an information descriptor:

  • The get_info() function template can be used to query general information that is available with any backend; and

  • The get_backend_info() function template can be used to query backend-specific information.

The information that can be queried with get_info() for a specific class is listed alongside the definition of that class. The information that can be queried with get_backend_info() is defined in the corresponding SYCL backend specification.

4.6. SYCL runtime classes

4.6.1. Device selection

Since a system can have several SYCL-compatible devices attached, it is useful to have a way to select a specific device or a set of devices to construct a specific object such as a device (see Section 4.6.4) or a queue (see Section 4.6.5), or perform some operations on a device subset.

Device selection is done either by already having a specific instance of a device (see Section 4.6.4) or by providing a device selector which is a ranking function that will give an integer ranking value to all the devices on the system.

4.6.1.1. Device selector

The interface for a device selector is any object that meets the C++ named requirement Callable, taking a parameter of type const device & and returning a value that is implicitly convertible to int.

At any point where the SYCL runtime needs to select a SYCL device using a device selector, the system queries all root devices from all SYCL backends in the system, calls the device selector on each device and selects the one which returns the highest score. If the highest value is strictly negative no device is selected.

In places where only one device has to be picked and the high score is obtained by more than one device, then one of the tied devices will be returned, but which one is not defined and may depend on enumeration order, for example, outside the control of the SYCL runtime.

Some predefined device selectors are provided by the system as described on Table 14 in a header file with some definition similar to the following:

Table 14. Standard device selectors included with all SYCL implementations
SYCL device selectors Description
default_selector_v

Select a SYCL device from any supported SYCL backend based on an implementation-defined heuristic. Since all implementations must support at least one device, this selector must always return a device.

Implementations may choose to return an emulated device (with aspect::emulated) as a fallback if there is no physical device available on the system.

gpu_selector_v

Select a SYCL device from any supported SYCL backend for which the device type is info::device_type::gpu. The SYCL class constructor using it must throw an exception with the errc::runtime error code if no device matching this requirement can be found.

accelerator_selector_v

Select a SYCL device from any supported SYCL backend for which the device type is info::device_type::accelerator. The SYCL class constructor using it must throw an exception with the errc::runtime error code if no device matching this requirement can be found.

cpu_selector_v

Select a SYCL device from any supported SYCL backend for which the device type is info::device_type::cpu. The SYCL class constructor using it must throw an exception with the errc::runtime error code if no device matching this requirement can be found.

__unspecified_callable__
aspect_selector(const std::vector<aspect>& aspectList,
                const std::vector<aspect>& denyList = {});

template <typename... AspectList>
__unspecified_callable__ aspect_selector(AspectList... aspectList);

template <aspect... AspectList> __unspecified_callable__ aspect_selector();

The free function aspect_selector has several overloads, each of which returns a selector object that selects a SYCL device from any supported SYCL backend which contains all the requested aspects, i.e. for the specific device dev and each aspect devAspect from aspectList dev.has(devAspect) equals true. If no aspects are passed in, the generated selector behaves like default_selector_v.

Required aspects can be passed in as a vector, as function arguments, or as template parameters, depending on the function overload. The function overload that takes aspectList as a vector takes another vector argument denyList where the user can specify all the aspects that have to be avoided, i.e. for the specific device dev and each aspect devAspect from denyList dev.has(devAspect) equals false.

The SYCL class constructor using the generated selector must throw an exception with the errc::runtime error code if no device matching this requirement can be found. There are multiple overloads of this function, please refer to [header:device-selector] for full definitions and to [example:aspect-selector] for examples.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
namespace sycl {

// Predefined device selectors
__unspecified__ default_selector_v;
__unspecified__ cpu_selector_v;
__unspecified__ gpu_selector_v;
__unspecified__ accelerator_selector_v;

// Predefined types for compatibility with old SYCL 1.2.1 device selectors
// Deprecated in SYCL 2020
using default_selector = __unspecified__;
using cpu_selector = __unspecified__;
using gpu_selector = __unspecified__;
using accelerator_selector = __unspecified__;

// Returns a selector that selects a device based on desired aspects
__unspecified_callable__
aspect_selector(const std::vector<aspect>& aspectList,
                const std::vector<aspect>& denyList = {});
template <class... AspectList>
__unspecified_callable__ aspect_selector(AspectList... aspectList);
template <aspect... AspectList> __unspecified_callable__ aspect_selector();

} // namespace sycl

Typical examples of default and user-provided device selectors could be:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
sycl::device my_gpu { sycl::gpu_selector_v };

sycl::queue my_accelerator { sycl::accelerator_selector_v };

int prefer_my_vendor(const sycl::device& d) {
  // Return 1 if the vendor name is "MyVendor" or 0 else.
  // 0 does not prevent another device to be picked as a second choice
  return d.get_info<info::device::vendor>() == "MyVendor";
}

// Get the preferred device or another one if not available
sycl::device preferred_device { prefer_my_vendor };

// This throws if there is no such device in the system
sycl::queue half_precision_controller {
  // Can use a lambda as a device ranking function.
  // Returns a negative number to fail in the case there is no such device
  [] (auto& d) { return d.has(sycl::aspect::fp16) ? 1 : -1; }
};

// To ease porting SYCL 1.2.1 code, there are types whose
// construction leads to the equivalent predefined device selector
sycl::queue my_old_style_gpu { sycl::gpu_selector {} };

Examples of using aspect_selector:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
using namespace sycl; // (optional) avoids need for "sycl::" before SYCL names

// Unrestrained selection, equivalent to default_selector_v
auto dev0 = device{aspect_selector()};

// Pass aspects in a vector
// Only accept CPUs that support half
auto dev1 = device{aspect_selector(std::vector{aspect::cpu, aspect::fp16})};

// Pass aspects without a vector
// Only accept GPUs that support half
auto dev2 = device{aspect_selector(aspect::gpu, aspect::fp16)};

// Pass aspects as compile-time parameters
// Only accept devices that can be debugged on host and support half
auto dev3 = device{aspect_selector<aspect::host_debuggable, aspect::fp16>()};

// Pass aspects in an allowlist and a denylist
// Only accept devices that support half and double floating point precision,
// but exclude emulated devices and devices of type "custom"
auto dev4 = device{aspect_selector(
   std::vector{aspect::fp16, aspect::fp64},
   std::vector{aspect::emulated, aspect::custom}
)};

In SYCL 1.2.1 the predefined device selectors were actually types that had to be instantiated to be used. Now they are just instances. To simplify porting code using the old type instantiations, a backward-compatible API is still provided, though deprecated, such as sycl::default_selector. The new predefined device selectors have their new names appended with "_v" to avoid conflicts, thus following the naming style used by traits in the C++ standard library. There is no requirement for the implementation to have for example sycl::gpu_selector_v being an instance of sycl::gpu_selector.

4.6.2. Platform class

The platform class encapsulates a single SYCL platform on which kernel functions may be executed. A platform must be associated with a single SYCL backend.

A platform also contains a set of devices that are associated with the same SYCL backend. A platform may contain no devices.

All member functions of the platform class are synchronous and errors are handled by throwing synchronous SYCL exceptions.

The execution environment for a SYCL application has a fixed number of platforms which does not vary as the application executes. The application can get a list of all these platforms via platform::get_platforms, and the order of the platform objects is the same each time the application calls that function. The platform class also provides constructors, but constructing a new platform instance merely creates a new object that is a copy of one of the objects returned by platform::get_platforms.

Each platform has an associated default context which contains all of the root devices in the platform. This default context does not have an asynchronous error handler. Applications can retrieve a copy of this default context object, for example, by constructing a queue. These copies follow the common reference semantics, as though they are all copies of an internal per-platform context object representing the platform’s default context.

The platform class provides the common reference semantics as defined in Section 4.5.2.

namespace sycl {
class platform {
 public:
  platform();

  template <typename DeviceSelector>
  explicit platform(const DeviceSelector& deviceSelector);

  /* -- common interface members -- */

  backend get_backend() const noexcept;

  std::vector<device>
      get_devices(info::device_type type = info::device_type::all) const;

  template <typename Param> typename Param::return_type get_info() const;

  template <typename Param>
  typename Param::return_type get_backend_info() const;

  bool has(aspect asp) const;

  bool has_extension(const std::string& extension) const; // Deprecated

  static std::vector<platform> get_platforms();
};
} // namespace sycl
4.6.2.1. Constructors
Default constructor
platform()

Effects: Constructs a platform object that is a copy of the platform which contains the device returned by default_selector_v.


Selector constructor
template <typename DeviceSelector>
explicit platform(const DeviceSelector& selector)

Constraints: The DeviceSelector must be a type that satisfies the requirements of a device selector as defined in Section 4.6.1.1.

Effects: The selector is called for every root device as described in Section 4.6.1.1. Constructs a platform object that is a copy of the platform which contains the device that is selected by selector.


4.6.2.2. Member functions
platform::get_backend
backend get_backend() const noexcept

Returns: The SYCL backend that is associated with this platform.


platform::get_info
template <typename Param>
typename Param::return_type get_info() const

Constraints: The Param must be an information descriptor for the platform class.

Each information descriptor specifies the return value and may also specify preconditions, exceptions that are thrown, etc. See Section 4.6.2.4 for the platform information descriptors that are defined by the core SYCL specification.


platform::get_backend_info
template <typename Param>
typename Param::return_type get_backend_info() const

Constraints: The Param must be a backend information descriptor for the platform class.

Throws: An exception with the errc::backend_mismatch error code if the backend that corresponds with Param is different from the backend that is associated with this platform.

Each information descriptor specifies the return value and may also specify preconditions, additional exceptions that are thrown, etc.


platform::has
bool has(aspect asp) const

Returns: The value true if all of the devices associated with this platform have the given aspect. Returns the value false if this platform does not contain any devices.


platform::has_extension
bool has_extension(const std::string& extension) const

Deprecated by SYCL 2020.

[Note: Use platform::has instead. — end note]

Returns: The value true if this platform supports the extension queried by the extension parameter. A platform only supports an extension if all associated devices support that extension. Returns false if this platform does not contain any devices.


platform::get_devices
std::vector<device>
get_devices(info::device_type type = info::device_type::all) const

Returns: A std::vector containing all of the root devices associated with this platform which have the device type specified by type.

[Note: Since the concept of a "host device" does not exist in SYCL 2020, if type is info::device_type::host this function will always return an empty vector.— end note]

Remarks: If type is info::device_type::all, the std::vector contains all root devices in this platform. If type is info::device_type::automatic and the platform is not empty, the std::vector contains a single root device corresponding to an implementation-defined default device for this platform. If the platform is empty, any call to this function returns an empty vector.


4.6.2.3. Static member functions
platform::get_platforms
static std::vector<platform> get_platforms()

Returns: A std::vector containing all of the platforms from all backends that are available in the system.


4.6.2.4. Information descriptors

This section describes the information descriptors that can be used as the Param template parameter to platform::get_info. When the description has a Returns, Throws, etc. paragraph, this indicates the value returned by or the exceptions thrown by the platform::get_info function.


info::platform::version
namespace sycl::info::platform {
struct version {
  using return_type = std::string;
};
} // namespace sycl::info::platform

Remarks: Template parameter to platform::get_info.

Returns: An implementation-defined platform version string.


info::platform::name
namespace sycl::info::platform {
struct name {
  using return_type = std::string;
};
} // namespace sycl::info::platform

Remarks: Template parameter to platform::get_info.

Returns: An implementation-defined name for this platform.


info::platform::vendor
namespace sycl::info::platform {
struct vendor {
  using return_type = std::string;
};
} // namespace sycl::info::platform

Remarks: Template parameter to platform::get_info.

Returns: An implementation-defined name for the vendor providing this platform.


info::platform::extensions
namespace sycl::info::platform {
struct extensions {
  using return_type = std::vector<std::string>;
};
} // namespace sycl::info::platform

Deprecated by SYCL 2020.

[Note: Use device::get_info() with info::device::aspects instead. — end note]

Remarks: Template parameter to platform::get_info.

Returns: The extensions supported by this platform. Returns an empty list if this platform does not contain any devices.


4.6.3. Context class

The context class represents a SYCL context. A context represents the runtime data structures and state required by a SYCL backend API to interact with a group of devices associated with a platform.

All member functions of the context class are synchronous and errors are handled by throwing synchronous SYCL exceptions.

All constructors of the context class construct an object that is associated with a particular SYCL backend, determined by the constructor parameters or, in the case of the default constructor, the SYCL device produced by the default_selector_v.

A context can optionally be constructed with an async_handler parameter. In this case the async_handler is used to report asynchronous exceptions, as described in Section 4.13.

The context class provides the common reference semantics as defined in Section 4.5.2.

namespace sycl {
class context {
 public:
  explicit context(const property_list& propList = {});

  explicit context(async_handler asyncHandler,
                   const property_list& propList = {});

  explicit context(const device& dev, const property_list& propList = {});

  explicit context(const device& dev, async_handler asyncHandler,
                   const property_list& propList = {});

  explicit context(const platform& plt, const property_list& propList = {});

  explicit context(const platform& plt, async_handler asyncHandler,
                   const property_list& propList = {});

  explicit context(const std::vector<device>& deviceList,
                   const property_list& propList = {});

  explicit context(const std::vector<device>& deviceList,
                   async_handler asyncHandler,
                   const property_list& propList = {});

  /* -- property interface members -- */

  /* -- common interface members -- */

  backend get_backend() const noexcept;

  platform get_platform() const;

  std::vector<device> get_devices() const;

  template <typename Param> typename Param::return_type get_info() const;

  template <typename Param>
  typename Param::return_type get_backend_info() const;
};
} // namespace sycl
4.6.3.1. Constructors

All context constructors take a parameter named propList which allows the application to pass zero or more properties. These properties may specify additional effects of the constructor and may also specify exceptions that the constructor throws. See Section 4.6.3.4 for the context properties that are defined by the core SYCL specification.


Default constructor
explicit context(const property_list& propList = {})

Effects: Constructs a context object using the device selected by default_selector_v. The context’s platform is the platform that contains this device. The context contains the selected device. The context may also contain other devices from the same platform. Whether this happens is implementation defined.


Constructor with async handler
explicit context(async_handler asyncHandler, const property_list& propList = {})

Effects: Constructs a context object using the device selected by default_selector_v. The context’s platform is the platform that contains this device. The context contains the selected device. The context may also contain other devices from the same platform. Whether this happens is implementation defined. The context has the asynchronous error handler asyncHandler.


Constructor with device
explicit context(const device& dev, const property_list& propList = {})

Effects: Constructs a context object that contains the device dev. The context’s platform is the platform that contains dev.


Constructor with device and async handler
explicit context(const device& dev, async_handler asyncHandler,
                 const property_list& propList = {})

Effects: Constructs a context object that contains the device dev. The context’s platform is the platform that contains dev. The context has the asynchronous error handler asyncHandler.


Constructor with platform
explicit context(const platform& plt, const property_list& propList = {})

Effects: Constructs a context object that contains all of the devices in the platform plt. The context’s platform is plt.

Throws: An exception with the errc::invalid error code if the platform plt contains no devices.


Constructor with platform and async handler
explicit context(const platform& plt, async_handler asyncHandler,
                 const property_list& propList = {})

Effects: Constructs a context object that contains all of the devices in the platform plt. The context’s platform is plt. The context has the asynchronous error handler asyncHandler.

Throws: An exception with the errc::invalid error code if the platform plt contains no devices.


Constructor with device list
explicit context(const std::vector<device>& deviceList, const property_list& propList = {})

Preconditions: All devices in deviceList must be associated with the same platform.

Effects: Constructs a context object that contains all of the devices in deviceList. The context’s platform is the platform that contains the devices in deviceList.

Throws: An exception with the errc::invalid error code if deviceList is empty.


Constructor with device list and async handler
explicit context(const std::vector<device>& deviceList, async_handler asyncHandler,
                 const property_list& propList = {})

Preconditions: All devices in deviceList must be associated with the same platform.

Effects: Constructs a context object that contains all of the devices in deviceList. The context’s platform is the platform that contains the devices in deviceList. The context has the asynchronous error handler asyncHandler.

Throws: An exception with the errc::invalid error code if deviceList is empty.


4.6.3.2. Member functions
context::get_backend
backend get_backend() const noexcept

Returns: The SYCL backend that is associated with this context.


context::get_platform
platform get_platform() const

Returns: The platform that is associated with this context.


context::get_devices
std::vector<device> get_devices() const

Returns: A std::vector containing all the devices that are associated with this context.


context::get_info
template <typename Param>
typename Param::return_type get_info() const

Constraints: The Param must be an information descriptor for the context class.

Each information descriptor specifies the return value and may also specify preconditions, exceptions that are thrown, etc. See Section 4.6.3.3 for the context information descriptors that are defined by the core SYCL specification.


context::get_backend_info
template <typename Param>
typename Param::return_type get_backend_info() const

Constraints: The Param must be a backend information descriptor for the context class.

Throws: An exception with the errc::backend_mismatch error code if the backend that corresponds with Param is different from the backend that is associated with this context.

Each information descriptor specifies the return value and may also specify preconditions, additional exceptions that are thrown, etc.


4.6.3.3. Information descriptors

This section describes the information descriptors that can be used as the Param template parameter to context::get_info. When the description has a Returns, Throws, etc. paragraph, this indicates the value returned by or the exceptions thrown by the context::get_info function.


info::context::platform
namespace sycl::info::context {
struct platform {
  using return_type = platform;
};
} // namespace sycl::info::context

Remarks: Template parameter to context::get_info.

Returns: The platform that is associated with this context.


info::context::devices
namespace sycl::info::context {
struct devices {
  using return_type = std::vector<device>;
};
} // namespace sycl::info::context

Remarks: Template parameter to context::get_info.

Returns: A std::vector containing all the devices that are associated with this context.


info::context::atomic_memory_order_capabilities
namespace sycl::info::context {
struct atomic_memory_order_capabilities {
  using return_type = std::vector<memory_order>;
};
} // namespace sycl::info::context

Remarks: Template parameter to context::get_info.

Returns: This query applies only to the capabilities of atomic operations that are applied to memory that can be concurrently accessed by multiple devices in the context. If these capabilities are not uniform across all devices in the context, the query reports only the capabilities that are common for all devices.

Returns the set of memory orders supported by these atomic operations. When a context returns a "stronger" memory order in this set, it must also return all "weaker" memory orders. (See Section 3.8.3.1 for a definition of "stronger" and "weaker" memory orders.) The memory orders memory_order::acquire, memory_order::release, and memory_order::acq_rel are all the same strength. If a context returns one of these, it must return them all.

At a minimum, each context must support memory_order::relaxed.


info::context::atomic_fence_order_capabilities
namespace sycl::info::context {
struct atomic_fence_order_capabilities {
  using return_type = std::vector<memory_order>;
};
} // namespace sycl::info::context

Remarks: Template parameter to context::get_info.

Returns: This query applies only to the capabilities of atomic_fence when applied to memory that can be concurrently accessed by multiple devices in the context. If these capabilities are not uniform across all devices in the context, the query reports only the capabilities that are common for all devices.

Returns the set of memory orders supported by these atomic_fence operations. When a context returns a "stronger" memory order in this set, it must also return all "weaker" memory orders. (See Section 3.8.3.1 for a definition of "stronger" and "weaker" memory orders.)

At a minimum, each context must support memory_order::relaxed, memory_order::acquire, memory_order::release, and memory_order::acq_rel.


info::context::atomic_memory_scope_capabilities
namespace sycl::info::context {
struct atomic_memory_scope_capabilities {
  using return_type = std::vector<memory_scope>;
};
} // namespace sycl::info::context

Remarks: Template parameter to context::get_info.

Returns: The set of memory scopes supported by atomic operations on all devices in the context, which is a subset of {memory_scope::sub_group, memory_scope::work_group, memory_scope::device, memory_scope::system}. When a context returns a "wider" memory scope in this set, it must also return all "narrower" memory scopes in this subset. (See Section 3.8.3.2 for a definition of "wider" and "narrower" scopes.) At a minimum, each context must support memory_scope::sub_group and memory_scope::work_group.


info::context::atomic_fence_scope_capabilities
namespace sycl::info::context {
struct atomic_fence_scope_capabilities {
  using return_type = std::vector<memory_scope>;
};
} // namespace sycl::info::context

Remarks: Template parameter to context::get_info.

Returns: The set of memory orderings supported by atomic_fence on all devices in the context. When a context returns a "wider" memory scope in this set, it must also return all "narrower" memory scopes. (See Section 3.8.3.2 for a definition of "wider" and "narrower" scopes.) At a minimum, each context must support memory_scope::work_item, memory_scope::sub_group, and memory_scope::work_group.


4.6.3.4. Properties

The core SYCL specification does not define any properties for the context constructors. The property_list constructor parameters are present for extensibility.

4.6.4. Device class

The device class represents a single SYCL device on which kernels can be executed.

All member functions of the device class are synchronous and errors are handled by throwing synchronous SYCL exceptions.

The execution environment for a SYCL application has a fixed number of root devices which does not vary as the application executes. The application can get a list of all these devices via device::get_devices, and the order of the device objects is the same each time the application calls that function (assuming the parameter to that function is the same for each call). The device class also provides constructors, but constructing a new device instance merely creates a new object that is a copy of one of the objects returned by device::get_devices.

A device can be partitioned into multiple devices, by calling the device::create_sub_devices member function template. The resulting device objects are considered sub-devices, and it is valid to partition these sub-devices further. The range of support for this feature is SYCL backend and device specific and can be queried for through device::get_info.

The device class provides the common reference semantics as defined in Section 4.5.2.

namespace sycl {

class device {
 public:
  device();

  template <typename DeviceSelector>
  explicit device(const DeviceSelector& deviceSelector);

  /* -- common interface members -- */

  backend get_backend() const noexcept;

  bool is_cpu() const;

  bool is_gpu() const;

  bool is_accelerator() const;

  platform get_platform() const;

  template <typename Param> typename Param::return_type get_info() const;

  template <typename Param>
  typename Param::return_type get_backend_info() const;

  bool has(aspect asp) const;

  bool has_extension(const std::string& extension) const; // Deprecated

  // Available only when Prop == info::partition_property::partition_equally
  template <info::partition_property Prop>
  std::vector<device> create_sub_devices(std::size_t count) const;

  // Available only when Prop == info::partition_property::partition_by_counts
  template <info::partition_property Prop>
  std::vector<device>
  create_sub_devices(const std::vector<std::size_t>& counts) const;

  // Available only when Prop ==
  // info::partition_property::partition_by_affinity_domain
  template <info::partition_property Prop>
  std::vector<device>
  create_sub_devices(info::partition_affinity_domain affinityDomain) const;

  static std::vector<device>
  get_devices(info::device_type type = info::device_type::all);
};
} // namespace sycl
4.6.4.1. Constructors
Default constructor
device()

Effects: Constructs a device object that is a copy of the device returned by default_selector_v.


Selector constructor
template <typename DeviceSelector>
explicit device(const DeviceSelector& selector)

Constraints: Available only when the DeviceSelector is a type that satisfies the requirements of a device selector as defined in Section 4.6.1.1.

Effects: The selector is called for every root device as described in Section 4.6.1.1. Constructs a device object that is a copy of the device selected by selector.


4.6.4.2. Member functions
device::get_backend
backend get_backend() const noexcept

Returns: The SYCL backend that is associated with this device.


device::get_platform
platform get_platform() const

Returns: The platform that is associated with this device.


device::is_cpu
bool is_cpu() const

Returns: The same value as has(aspect::cpu). See Section 4.6.4.5.


device::is_gpu
bool is_gpu() const

Returns: The same value as has(aspect::gpu). See Section 4.6.4.5.


device::is_accelerator
bool is_accelerator() const

Returns: The same value as has(aspect::accelerator). See Section 4.6.4.5.


device::get_info
template <typename Param>
typename Param::return_type get_info() const

Constraints: Available only when Param is an information descriptor for the device class.

Each information descriptor specifies the return value and may also specify preconditions, exceptions that are thrown, etc. See Section 4.6.4.4 for the device information descriptors that are defined by the core SYCL specification.


device::get_backend_info
template <typename Param>
typename Param::return_type get_backend_info() const

Constraints: Available only when Param is a backend information descriptor for the device class.

Throws: An exception with the errc::backend_mismatch error code if the backend that corresponds with Param is different from the backend that is associated with this device.

Each information descriptor specifies the return value and may also specify preconditions, additional exceptions that are thrown, etc.


device::has
bool has(aspect asp) const

Returns: The value true if this device has the given aspect. Applications can use this member function to determine which optional features this device supports (if any).


device::has_extension
bool has_extension(const std::string& extension) const

Deprecated by SYCL 2020.

[Note: Use device::has instead. — end note]

Returns: The value true if this device supports the extension queried by the extension parameter.


device::create_sub_devices (partition equally)
template <info::partition_property Prop>
std::vector<device> create_sub_devices(std::size_t count) const

Constraints: Available only when Prop is info::partition_property::partition_equally.

Returns: A std::vector of sub-devices partitioned from this device object based on the count parameter. The returned vector contains as many sub-devices as can be created such that each sub-device contains count compute units. If the device’s total number of compute units (as returned by info::device::max_compute_units) is not evenly divided by count, then the remaining compute units are not included in any of the sub-devices.

Throws:

  • An exception with the errc::feature_not_supported error code if this device does not support info::partition_property::partition_equally.

  • An exception with the errc::invalid error code if count exceeds the total number of compute units in the device.


device::create_sub_devices (partition by counts)
template <info::partition_property Prop>
std::vector<device> create_sub_devices(const std::vector<std::size_t>& counts) const

Constraints: Available only when Prop is info::partition_property::partition_by_counts.

Returns: A std::vector of sub-devices partitioned from this device object based on the counts parameter. For each non-zero value M in the counts vector, a sub-device with M compute units is created.

Throws:


device::create_sub_devices (partition by affinity domain)
template <info::partition_property Prop>
std::vector<device>
create_sub_devices(info::partition_affinity_domain domain) const

Constraints: Available only when Prop is info::partition_property::partition_by_affinity_domain.

Returns: A std::vector of sub-devices partitioned from this device object based on the domain parameter, which must be one of the following values:

Throws:


4.6.4.3. Static member functions
device::get_devices
static std::vector<device>
get_devices(info::device_type type = info::device_type::all)

Returns: A std::vector containing all the root devices from all SYCL backends available in the system which have the device type type.

[Note: Since the concept of a "host device" does not exist in SYCL 2020, if type is info::device_type::host this function will always return an empty vector.— end note]

Remarks: If type is info::device_type::all, the std::vector contains all root devices in the system. If type is info::device_type::automatic, the std::vector contains one root device from each non-empty platform, corresponding to the device returned by platform::get_devices(info::device_type::automatic).


4.6.4.4. Information descriptors

This section describes the information descriptors that can be used as the Param template parameter to device::get_info. When the description has a Returns, Throws, etc. paragraph, this indicates the value returned by or the exceptions thrown by the device::get_info function.


info::device::device_type
namespace sycl::info::device {
struct device_type {
  using return_type = info::device_type;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The device type associated with the device. May not return info::device_type::all or info::device_type::automatic.


info::device::vendor_id
namespace sycl::info::device {
struct vendor_id {
  using return_type = std::uint32_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: A unique vendor device identifier.


info::device::max_compute_units
namespace sycl::info::device {
struct max_compute_units {
  using return_type = std::uint32_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The number of parallel compute units available to the device. The minimum value is 1.


info::device::max_work_item_dimensions
namespace sycl::info::device {
struct max_work_item_dimensions {
  using return_type = std::uint32_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The maximum dimensions that specify the global and local work-item IDs used by the data parallel execution model. The minimum value is 3 if this device is not of device type info::device_type::custom.


info::device::max_work_item_sizes
namespace sycl::info::device {
template<int Dimensions = 3>
struct max_work_item_sizes<Dimensions> {
  using return_type = range<Dimensions>;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Constraints: Available only when Dimensions is 1, 2, or 3.

Returns: The maximum number of work-items that are permitted in a work-group for a kernel running in an index space of Dimensions dimensions. When the device type is not info::device_type::custom, the minimum value returned from this query is: (1) when Dimensions is 1, (1, 1) when Dimensions is 2, and (1, 1, 1) when Dimensions is 3.


info::device::max_work_group_size
namespace sycl::info::device {
struct max_work_group_size {
  using return_type = std::size_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The maximum number of work-items that this device is capable of executing in a work-group. The minimum value is 1. This value is an upper limit and will not necessarily maximize performance. The maximum number of work-items in a work-group depends on the kernel and the implementation. Use info::kernel_device_specific::work_group_size to query this limit.


info::device::max_num_sub_groups
namespace sycl::info::device {
struct max_num_sub_groups {
  using return_type = std::uint32_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The maximum number of sub-groups that this device is capable of executing in a work-group. The minimum value is 1. The maximum number of sub-groups in a work-group depends on the kernel and the implementation. Use info::kernel_device_specific::max_num_sub_groups to query this limit.

[Note: The largest work-group size supported by a device is likely to be the product of max_num_sub_groups and the largest supported sub-group size.— end note]


info::device::sub_group_sizes
namespace sycl::info::device {
struct sub_group_sizes {
  using return_type = std::vector<std::size_t>;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: A std::vector of std::size_t containing the set of sub-group sizes supported by the device.


info::device::preferred_vector_width
namespace sycl::info::device {
struct preferred_vector_width_char {
  using return_type = std::uint32_t;
};
struct preferred_vector_width_short {
  using return_type = std::uint32_t;
};
struct preferred_vector_width_int {
  using return_type = std::uint32_t;
};
struct preferred_vector_width_long {
  using return_type = std::uint32_t;
};
struct preferred_vector_width_long_long {
  using return_type = std::uint32_t;
};
struct preferred_vector_width_float {
  using return_type = std::uint32_t;
};
struct preferred_vector_width_double {
  using return_type = std::uint32_t;
};
struct preferred_vector_width_half {
  using return_type = std::uint32_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The preferred native vector width size for built-in scalar types that can be put into vectors. The vector width is defined as the number of scalar elements that can be stored in the vector. Must return 0 for info::device::preferred_vector_width_double if the device does not have aspect::fp64 and must return 0 for info::device::preferred_vector_width_half if the device does not have aspect::fp16.


info::device::native_vector_width
namespace sycl::info::device {
struct native_vector_width_char {
  using return_type = std::uint32_t;
};
struct native_vector_width_short {
  using return_type = std::uint32_t;
};
struct native_vector_width_int {
  using return_type = std::uint32_t;
};
struct native_vector_width_long {
  using return_type = std::uint32_t;
};
struct native_vector_width_long_long {
  using return_type = std::uint32_t;
};
struct native_vector_width_float {
  using return_type = std::uint32_t;
};
struct native_vector_width_double {
  using return_type = std::uint32_t;
};
struct native_vector_width_half {
  using return_type = std::uint32_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The native ISA vector width. The vector width is defined as the number of scalar elements that can be stored in the vector. Must return 0 for info::device::native_vector_width_double if the device does not have aspect::fp64 and must return 0 for info::device::native_vector_width_half if the device does not have aspect::fp16.


info::device::max_clock_frequency
namespace sycl::info::device {
struct max_clock_frequency {
  using return_type = std::uint32_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The maximum configured clock frequency of this device in MHz.


info::device::address_bits
namespace sycl::info::device {
struct address_bits {
  using return_type = std::uint32_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The default compute device address space size in bits. Must return either 32 or 64.


info::device::max_mem_alloc_size
namespace sycl::info::device {
struct max_mem_alloc_size {
  using return_type = std::uint64_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The maximum size of memory object allocation in bytes.


info::device::image_support
namespace sycl::info::device {
struct image_support {
  using return_type = bool;
};
} // namespace sycl::info::device

Deprecated by SYCL 2020.

Remarks: Template parameter to device::get_info.

Returns: The same value as device::has(aspect::image).


info::device::max_read_image_args
namespace sycl::info::device {
struct max_read_image_args {
  using return_type = std::uint32_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The maximum number of simultaneous image objects that can be read from by a kernel. The minimum value is 128 if the device has aspect::image.


info::device::max_write_image_args
namespace sycl::info::device {
struct max_write_image_args {
  using return_type = std::uint32_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The maximum number of simultaneous image objects that can be written to by a kernel. The minimum value is 8 if the device has aspect::image.


info::device::image2d_max_width
namespace sycl::info::device {
struct image2d_max_width {
  using return_type = std::size_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The maximum width of a 2D image or 1D image in pixels. The minimum value is 8192 if the device has aspect::image.


info::device::image2d_max_height
namespace sycl::info::device {
struct image2d_max_height {
  using return_type = std::size_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The maximum height of a 2D image in pixels. The minimum value is 8192 if the device has aspect::image.


info::device::image3d_max_width
namespace sycl::info::device {
struct image3d_max_width {
  using return_type = std::size_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The maximum width of a 3D image in pixels. The minimum value is 2048 if the device has aspect::image.


info::device::image3d_max_height
namespace sycl::info::device {
struct image3d_max_height {
  using return_type = std::size_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The maximum height of a 3D image in pixels. The minimum value is 2048 if the device has aspect::image.


info::device::image3d_max_depth
namespace sycl::info::device {
struct image3d_max_depth {
  using return_type = std::size_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The maximum depth of a 3D image in pixels. The minimum value is 2048 if the device has aspect::image.


info::device::image_max_buffer_size
namespace sycl::info::device {
struct image_max_buffer_size {
  using return_type = std::size_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The number of pixels for a 1D image created from a buffer object. The minimum value is 65536 if the device has aspect::image. Note that this information is intended for OpenCL interoperability only as this feature is not supported in SYCL.


info::device::max_samplers
namespace sycl::info::device {
struct max_samplers {
  using return_type = std::uint32_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The maximum number of samplers that can be used in a kernel. The minimum value is 16 if the device has aspect::image.


info::device::max_parameter_size
namespace sycl::info::device {
struct max_parameter_size {
  using return_type = std::size_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The maximum size in bytes of the arguments that can be passed to a kernel. The minimum value is 1024 if this device is not of device type info::device_type::custom. For this minimum value, only a maximum of 128 arguments can be passed to a kernel.


info::device::mem_base_addr_align
namespace sycl::info::device {
struct mem_base_addr_align {
  using return_type = std::uint32_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The minimum value in bits of the largest supported SYCL built-in data type if this device is not of device type info::device_type::custom.


info::device::half_fp_config
namespace sycl::info::device {
struct half_fp_config {
  using return_type = std::vector<info::fp_config>;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: A std::vector of info::fp_config values describing the half precision floating-point capability of this device. The std::vector may contain zero or more of the following values:

If half precision is supported by this device (i.e. the device has aspect::fp16) there is no minimum floating-point capability. If half support is not supported the returned std::vector must be empty.


info::device::single_fp_config
namespace sycl::info::device {
struct single_fp_config {
  using return_type = std::vector<info::fp_config>;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: A std::vector of info::fp_config values describing the single precision floating-point capability of this device. The std::vector must contain one or more of the following values:

If this device is not of type info::device_type::custom then the minimum floating-point capability must be: info::fp_config::round_to_nearest and info::fp_config::inf_nan.


info::device::double_fp_config
namespace sycl::info::device {
struct double_fp_config {
  using return_type = std::vector<info::fp_config>;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: A std::vector of info::fp_config values describing the double precision floating-point capability of this device. The std::vector may contain zero or more of the following values:

If double precision is supported by this device (i.e. the device has aspect::fp64) and this device is not of type info::device_type::custom then the minimum floating-point capability must be: info::fp_config::fma, info::fp_config::round_to_nearest, info::fp_config::round_to_zero, info::fp_config::round_to_inf, info::fp_config::inf_nan and info::fp_config::denorm. If double support is not supported the returned std::vector must be empty.


info::device::global_mem_cache_type
namespace sycl::info::device {
struct global_mem_cache_type {
  using return_type = info::global_mem_cache_type;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The type of global memory cache supported.


info::device::global_mem_cache_line_size
namespace sycl::info::device {
struct global_mem_cache_line_size {
  using return_type = std::uint32_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The size of global memory cache line in bytes.


info::device::global_mem_cache_size
namespace sycl::info::device {
struct global_mem_cache_size {
  using return_type = std::uint64_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The size of global memory cache in bytes.


info::device::global_mem_size
namespace sycl::info::device {
struct global_mem_size {
  using return_type = std::uint64_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The size of global device memory in bytes.


info::device::max_constant_buffer_size
namespace sycl::info::device {
struct max_constant_buffer_size {
  using return_type = std::uint64_t;
};
} // namespace sycl::info::device

Deprecated by SYCL 2020.

Remarks: Template parameter to device::get_info.

Returns: The maximum size in bytes of a constant buffer allocation. The minimum value is 64 KB if this device is not of type info::device_type::custom.


info::device::max_constant_args
namespace sycl::info::device {
struct max_constant_args {
  using return_type = std::uint32_t;
};
} // namespace sycl::info::device

Deprecated by SYCL 2020.

Remarks: Template parameter to device::get_info.

Returns: The maximum number of constant arguments that can be declared in a kernel. The minimum value is 8 if this device is not of type info::device_type::custom.


info::device::local_mem_type
namespace sycl::info::device {
struct local_mem_type {
  using return_type = info::local_mem_type;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The type of local memory supported. This can be info::local_mem_type::local implying dedicated local memory storage such as SRAM, or info::local_mem_type::global. If this device is of type info::device_type::custom this can also be info::local_mem_type::none, indicating local memory is not supported.


info::device::local_mem_size
namespace sycl::info::device {
struct local_mem_size {
  using return_type = std::uint64_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The size of local memory arena in bytes. The minimum value is 32 KB if this device is not of type info::device_type::custom.


info::device::error_correction_support
namespace sycl::info::device {
struct error_correction_support {
  using return_type = bool;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The value true if the device implements error correction for all accesses to compute device memory (global and constant). Returns false if the device does not implement such error correction.


info::device::host_unified_memory
namespace sycl::info::device {
struct host_unified_memory {
  using return_type = bool;
};
} // namespace sycl::info::device

Deprecated by SYCL 2020.

[Note: Use device::has with one of the aspect::usm_* aspects instead. — end note]

Remarks: Template parameter to device::get_info.

Returns: The value true if the device and the host have a unified memory subsystem and returns false otherwise.


info::device::atomic_memory_order_capabilities
namespace sycl::info::device {
struct atomic_memory_order_capabilities {
  using return_type = std::vector<memory_order>;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The set of memory orders supported by atomic operations on this device. When a device returns a "stronger" memory order in this set, it must also return all "weaker" memory orders. (See Section 3.8.3.1 for a definition of "stronger" and "weaker" memory orders.) The memory orders memory_order::acquire, memory_order::release, and memory_order::acq_rel are all the same strength. If a device returns one of these, it must return them all.

At a minimum, each device must support memory_order::relaxed.


info::device::atomic_fence_order_capabilities
namespace sycl::info::device {
struct atomic_fence_order_capabilities {
  using return_type = std::vector<memory_order>;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The set of memory orders supported by atomic_fence on this device. When a device returns a "stronger" memory order in this set, it must also return all "weaker" memory orders. (See Section 3.8.3.1 for a definition of "stronger" and "weaker" memory orders.) At a minimum, each device must support memory_order::relaxed, memory_order::acquire, memory_order::release, and memory_order::acq_rel.


info::device::atomic_memory_scope_capabilities
namespace sycl::info::device {
struct atomic_memory_scope_capabilities {
  using return_type = std::vector<memory_scope>;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The set of memory scopes supported by atomic operations on this device, which is a subset of {memory_scope::sub_group, memory_scope::work_group, memory_scope::device, memory_scope::system}. When a device returns a "wider" memory scope in this set, it must also return all "narrower" memory scopes in this subset. (See Section 3.8.3.2 for a definition of "wider" and "narrower" scopes.) At a minimum, each device must support memory_scope::sub_group and memory_scope::work_group.


info::device::atomic_fence_scope_capabilities
namespace sycl::info::device {
struct atomic_fence_scope_capabilities {
  using return_type = std::vector<memory_scope>;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The set of memory scopes supported by atomic_fence on this device. When a device returns a "wider" memory scope in this set, it must also return all "narrower" memory scopes. (See Section 3.8.3.2 for a definition of "wider" and "narrower" scopes.) At a minimum, each device must support memory_scope::work_item, memory_scope::sub_group, and memory_scope::work_group.


info::device::profiling_timer_resolution
namespace sycl::info::device {
struct profiling_timer_resolution {
  using return_type = std::size_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The resolution of device timer in nanoseconds.


info::device::is_endian_little
namespace sycl::info::device {
struct is_endian_little {
  using return_type = bool;
};
} // namespace sycl::info::device

Deprecated by SYCL 2020.

[Note: Check the byte order of the host system instead. The host and device are required to have the same byte order. — end note]

Remarks: Template parameter to device::get_info.

Returns: The value true if this device is a little endian device and returns false otherwise.


info::device::is_available
namespace sycl::info::device {
struct is_available {
  using return_type = bool;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The value true if the device is available and false if the device is not available. A device is considered to be available if the device can be expected to successfully execute commands enqueued to the device. The conditions that lead to a device being considered available or not available are implementation-defined.


info::device::is_compiler_available
namespace sycl::info::device {
struct is_compiler_available {
  using return_type = bool;
};
} // namespace sycl::info::device

Deprecated by SYCL 2020.

Remarks: Template parameter to device::get_info.

Returns: The same value as device::has(aspect::online_compiler).


info::device::is_linker_available
namespace sycl::info::device {
struct is_linker_available {
  using return_type = bool;
};
} // namespace sycl::info::device

Deprecated by SYCL 2020.

Remarks: Template parameter to device::get_info.

Returns: The same value as device::has(aspect::online_linker).


info::device::execution_capabilities
namespace sycl::info::device {
struct execution_capabilities {
  using return_type = std::vector<info::execution_capability>;
};
} // namespace sycl::info::device

Deprecated by SYCL 2020.

Remarks: Template parameter to device::get_info.

Returns: Only supported when the backend of this device is OpenCL (see Appendix C). Returns a std::vector of the info::execution_capability values describing the supported execution capabilities. Unless the device type is info::device_type::custom, the returned vector will always include info::execution_capability::exec_kernel.

Throws: An exception with the errc::invalid error code if the backend of this device is not OpenCL.


info::device::queue_profiling
namespace sycl::info::device {
struct queue_profiling {
  using return_type = bool;
};
} // namespace sycl::info::device

Deprecated by SYCL 2020.

Remarks: Template parameter to device::get_info.

Returns: The same value as device::has(aspect::queue_profiling).


info::device::built_in_kernel_ids
namespace sycl::info::device {
struct built_in_kernel_ids {
  using return_type = std::vector<kernel_id>;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: A std::vector of identifiers for the built-in kernels supported by this device.


info::device::built_in_kernels
namespace sycl::info::device {
struct built_in_kernels {
  using return_type = std::vector<std::string>;
};
} // namespace sycl::info::device

Deprecated by SYCL 2020.

[Note: Use info::device::built_in_kernel_ids instead. — end note]

Remarks: Template parameter to device::get_info.

Returns: A std::vector of built-in OpenCL kernels supported by this device.


info::device::platform
namespace sycl::info::device {
struct platform {
  using return_type = platform;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The platform that is associated with this device.


info::device::name
namespace sycl::info::device {
struct name {
  using return_type = std::string;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: An implementation-defined name for this device.


info::device::vendor
namespace sycl::info::device {
struct vendor {
  using return_type = std::string;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: An implementation-defined name for the vendor providing this device.


info::device::driver_version
namespace sycl::info::device {
struct driver_version {
  using return_type = std::string;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: An implementation-defined name describing the version of the underlying software driver for this device.


info::device::profile
namespace sycl::info::device {
struct profile {
  using return_type = std::string;
};
} // namespace sycl::info::device

Deprecated by SYCL 2020.

Remarks: Template parameter to device::get_info.

Returns: Only supported when the backend of this device is OpenCL (see Appendix C). The value returned can be one of the following strings:

  • FULL_PROFILE - if the device supports the OpenCL specification (functionality defined as part of the core specification and does not require any extensions to be supported).

  • EMBEDDED_PROFILE - if the device supports the OpenCL embedded profile.

Throws: An exception with the errc::invalid error code if the backend of this device is not OpenCL.


info::device::version
namespace sycl::info::device {
struct version {
  using return_type = std::string;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: A backend-defined device version.


info::device::backend_version
namespace sycl::info::device {
struct backend_version {
  using return_type = std::string;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: A string describing the version of the SYCL backend associated with this device. The value returned from this query is defined by the backend interoperation specification that corresponds to this device’s backend.


info::device::aspects
namespace sycl::info::device {
struct aspects {
  using return_type = std::vector<aspect>;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: A std::vector of aspect values supported by this device.


info::device::extensions
namespace sycl::info::device {
struct extensions {
  using return_type = std::vector<std::string>;
};
} // namespace sycl::info::device

Deprecated by SYCL 2020.

[Note: Use info::device::aspects instead. — end note]

Remarks: Template parameter to device::get_info.

Returns: A std::vector of extension names (the extension names do not contain any spaces) supported by this device. The extension names returned can be vendor supported extension names and one or more of the following Khronos approved extension names:

  • cl_khr_int64_base_atomics

  • cl_khr_int64_extended_atomics

  • cl_khr_3d_image_writes

  • cl_khr_fp16

  • cl_khr_gl_sharing

  • cl_khr_gl_event

  • cl_khr_d3d10_sharing

  • cl_khr_dx9_media_sharing

  • cl_khr_d3d11_sharing

  • cl_khr_depth_images

  • cl_khr_gl_depth_images

  • cl_khr_gl_msaa_sharing

  • cl_khr_image2d_from_buffer

  • cl_khr_initialize_memory

  • cl_khr_context_abort

  • cl_khr_spir

If the backend associated with this device is OpenCL, then following approved Khronos extension names must be returned by all device that support OpenCL C 1.2:

  • cl_khr_global_int32_base_atomics

  • cl_khr_global_int32_extended_atomics

  • cl_khr_local_int32_base_atomics

  • cl_khr_local_int32_extended_atomics

  • cl_khr_byte_addressable_store

  • cl_khr_fp64 (for backward compatibility if double precision is supported)

Please refer to the OpenCL 1.2 Extension Specification for a detailed description of these extensions.


info::device::printf_buffer_size
namespace sycl::info::device {
struct printf_buffer_size {
  using return_type = std::size_t;
};
} // namespace sycl::info::device

Deprecated by SYCL 2020.

Remarks: Template parameter to device::get_info.

Returns: The maximum size of the internal buffer that holds the output of printf calls from a kernel. The minimum value is 1 MB if info::device::profile returns true for this device.


info::device::preferred_interop_user_sync
namespace sycl::info::device {
struct preferred_interop_user_sync {
  using return_type = bool;
};
} // namespace sycl::info::device

Deprecated by SYCL 2020.

Remarks: Template parameter to device::get_info.

Returns: Only supported when the backend of this device is OpenCL (see Appendix C). Returns true if the preference for this device is for the user to be responsible for synchronization, when sharing memory objects between OpenCL and other APIs such as DirectX, false if the device/implementation has a performant path for performing synchronization of memory object shared between OpenCL and other APIs such as DirectX.

Throws: An exception with the errc::invalid error code if the backend of this device is not OpenCL.


info::device::parent_device
namespace sycl::info::device {
struct parent_device {
  using return_type = device;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The parent device to which this sub-device is a child if this is a sub-device.

Throws: An exception with the errc::invalid error code if this device is not a sub-device.


info::device::partition_max_sub_devices
namespace sycl::info::device {
struct partition_max_sub_devices {
  using return_type = std::uint32_t;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The maximum number of sub-devices that can be created when this device is partitioned. The value returned cannot exceed the value returned by info::device::max_compute_units.


info::device::partition_properties
namespace sycl::info::device {
struct partition_properties {
  using return_type = std::vector<info::partition_property>;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: A std::vector of the partition properties supported by this device. An element is returned in this vector only if the device can be partitioned into at least two sub-devices along that partition property.


info::device::partition_affinity_domains
namespace sycl::info::device {
struct partition_affinity_domains {
  using return_type = std::vector<info::partition_affinity_domain>;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: A std::vector of the partition affinity domains supported by this device when partitioning with info::partition_property::partition_by_affinity_domain. An element is returned in this vector only if the device can be partitioned into at least two sub-devices along that affinity domain.


info::device::partition_type_property
namespace sycl::info::device {
struct partition_type_property {
  using return_type = info::partition_property;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The partition property of this device. If this device is not a sub-device then the return value is info::partition_property::no_partition, otherwise it is one of the following values:


info::device::partition_type_affinity_domain
namespace sycl::info::device {
struct partition_type_affinity_domain {
  using return_type = info::partition_affinity_domain;
};
} // namespace sycl::info::device

Remarks: Template parameter to device::get_info.

Returns: The partition affinity domain of this device. If this device is not a sub-device or the sub-device was not partitioned with info::partition_property::partition_by_affinity_domain then the return value is info::partition_affinity_domain::not_applicable, otherwise it is one of the following values:


4.6.4.5. Aspects

Every device has an associated set of aspects which identify characteristics of the device. Aspects are defined via the aspect enumeration:

namespace sycl {

enum class aspect : /* unspecified */ {
  cpu,
  gpu,
  accelerator,
  custom,
  emulated,
  host_debuggable,
  fp16,
  fp64,
  atomic64,
  image,
  online_compiler,
  online_linker,
  queue_profiling,
  usm_device_allocations,
  usm_host_allocations,
  usm_atomic_host_allocations,
  usm_shared_allocations,
  usm_atomic_shared_allocations,
  usm_system_allocations
};

} // namespace sycl

Applications can query the aspects of a device via device::has in order to determine whether the device supports any optional features. The following list describes the aspects that are defined in the core SYCL specification and tells which optional features correspond to each. Backends and extensions may provide additional aspects and additional optional device features. If so, the SYCL backend specification document or the extension document describes them.


aspect::cpu

Indicates that the implementation identifies this device as a CPU that has device type info::device_type::cpu.

[Note: A device with this aspect will typically share some or all of the execution resources available to the host C++ application. — end note]


aspect::gpu

Indicates that the implementation identifies this device as a GPU that has device type info::device_type::gpu.

[Note: A device with this aspect may have additional capabilities for accelerating graphics operations, via SYCL image functionality and/or interoperability with graphics APIs. — end note]


aspect::accelerator

Indicates that the implementation identifies this device as an accelerator that has device type info::device_type::accelerator.

[Note: A device with this aspect will typically be a dedicated accelerator device, with a peripheral interconnect for communication. — end note]


aspect::custom

Indicates that this device is a custom accelerator that exposes only fixed functionality, and has device type info::device_type::custom.

A device with this aspect does not support execution of arbitrary kernels, and can only execute pre-defined kernels (see Section 3.9.7).


aspect::emulated

Indicates that the device is somehow emulated.

A device with this aspect is not intended for performance, and instead will generally have another purpose such as emulation or profiling. The precise definition of this aspect is left open to the SYCL implementation.

[Note: As an example, a vendor might support both a hardware FPGA device and a software emulated FPGA, where the emulated FPGA has all the same features as the hardware one but runs more slowly and can provide additional profiling or diagnostic information. In such a case, an application’s device selector can use aspect::emulated to distinguish the two. — end note]


aspect::host_debuggable

Indicates that kernels running on this device can be debugged using standard debuggers that are normally available on the host system where the SYCL implementation resides. The precise definition of this aspect is left open to the SYCL implementation.


aspect::fp16

Indicates that kernels submitted to the device may use the sycl::half data type.


aspect::fp64

Indicates that kernels submitted to the device may use the double data type.


aspect::atomic64

Indicates that kernels submitted to the device may perform 64-bit atomic operations.


aspect::image

Indicates that the device supports images.


aspect::online_compiler

Indicates that the device supports online compilation of device code. Devices that have this aspect support the build and compile functions defined in Section 4.11.11.


aspect::online_linker

Indicates that the device supports online linking of device code. Devices that have this aspect support the link functions defined in Section 4.11.11. All devices that have this aspect also have aspect::online_compiler.


aspect::queue_profiling

Indicates that the device supports queue profiling via property::queue::enable_profiling.


aspect::usm_device_allocations

Indicates that the device supports explicit USM allocations as described in Section 4.8.


aspect::usm_host_allocations

Indicates that the device can access USM memory allocated via usm::alloc::host. Concurrent access and atomic modification of a host allocation is only supported if aspect::usm_atomic_host_allocations is also supported. (See Section 4.8.)


aspect::usm_atomic_host_allocations

Indicates that the device supports USM memory allocated via usm::alloc::host. The host and this device may concurrently access and atomically modify host allocations. (See Section 4.8.)


aspect::usm_shared_allocations

Indicates that the device supports USM memory allocated via usm::alloc::shared on the same device. Concurrent access and atomic modification of a shared allocation is only supported if aspect::usm_atomic_shared_allocations is also supported. (See Section 4.8.)


aspect::usm_atomic_shared_allocations

Indicates that the device supports USM memory allocated via usm::alloc::shared. The host and other devices in the same context that also support this capability may concurrently access and atomically modify shared allocations. The allocation is free to migrate between the host and the appropriate devices. (See Section 4.8.)


aspect::usm_system_allocations

Indicates that the system allocator may be used instead of SYCL USM allocation mechanisms for usm::alloc::shared allocations on this device. (See Section 4.8.)


4.6.4.6. Aspect traits

The implementation also provides two traits that the application can use to query aspects at compilation time. The traits any_device_has<aspect> and all_devices_have<aspect> are set according to the collection of devices D that can possibly execute device code, as determined by the compilation environment. The trait any_device_has<aspect> inherits from std::true_type only if at least one device in D has the specified aspect. The trait all_devices_have<aspect> inherits from std::true_type only if all devices in D have the specified aspect.

namespace sycl {

template <aspect Aspect> struct any_device_has;
template <aspect Aspect> struct all_devices_have;

template <aspect A>
inline constexpr bool any_device_has_v = any_device_has<A>::value;
template <aspect A>
inline constexpr bool all_devices_have_v = all_devices_have<A>::value;

} // namespace sycl

Applications can use these traits to reduce their code size. The following example demonstrates one way to use these traits to avoid instantiating a templated kernel for device features that are not supported by any device.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
#include <sycl/sycl.hpp>
using namespace sycl;  // (optional) avoids need for "sycl::" before SYCL names

constexpr int N = 512;

template <bool HasFp16>
class MyKernel {
 public:
  void operator()(id<1> i) {
    if constexpr (HasFp16) {
      // Algorithm using sycl::half type
    } else {
      // Fall back code for devices that don't support sycl::half
    }
  }
};

int main() {
  queue myQueue;
  myQueue.submit([&](handler& cgh) {
    device dev = myQueue.get_device();
    if (dev.has(aspect::fp16)) {
      cgh.parallel_for(range{N}, MyKernel<any_device_has_v<aspect::fp16>>{});
    } else {
      cgh.parallel_for(range{N}, MyKernel<all_devices_have_v<aspect::fp16>>{});
    }
  });

  myQueue.wait();
}

The kernel function MyKernel is templated to use a different algorithm depending on whether the device has the aspect aspect::fp16, and the call to dev.has() chooses the kernel function instantiation that matches the device’s capabilities. However, the use of any_device_has_v and all_devices_have_v entirely avoid useless instantiations of the kernel function. For example, when the compilation environment does not support any devices with aspect::fp16, any_device_has_v<aspect::fp16> is false, and the kernel function is never instantiated with support for the sycl::half type.

[Note: Like any trait, the definitions of any_device_has and all_devices_have are uniform across all parts of a SYCL application. If an implementation uses SMCP, all compiler passes define a particular aspect’s specialization of the traits the same way, regardless of whether that compiler pass' device supports the aspect. Thus, any_device_has and all_devices_have cannot be used to determine whether any particular device supports an aspect. Instead, applications must use device::has or platform::has for this. — end note]

[Note: An implementation could choose to provide command line options which affect the set of devices that it supports. If so, those command line options would also affect these traits. For example, if an implementation provides a command line option that disables aspect::accelerator devices, the trait any_device_has<aspect::accelerator> would inherit from std::false_type when that command line option was specified. — end note]

[Note: These traits only reflect the supported devices at the time the SYCL application is compiled. It’s possible that unsupported devices are still visible to the application when it runs. However, if a device D is not supported when the application is compiled, the application will not be able to submit kernels to that device D. — end note]

4.6.4.7. Other enumerations
4.6.4.7.1. Device type
namespace sycl::info {
enum class device_type : /* unspecified */ {
  cpu,
  gpu,
  accelerator,
  custom,
  automatic,
  host, // Deprecated by SYCL 2020
  all
};
} // namespace sycl::info
4.6.4.7.2. Partition property
namespace sycl::info {
enum class partition_property : /* unspecified */ {
  no_partition,
  partition_equally,
  partition_by_counts,
  partition_by_affinity_domain
};
} // namespace sycl::info
4.6.4.7.3. Partition affinity domain
namespace sycl::info {
enum class partition_affinity_domain : /* unspecified */ {
  not_applicable,
  numa,
  L4_cache,
  L3_cache,
  L2_cache,
  L1_cache,
  next_partitionable
};
} // namespace sycl::info
4.6.4.7.4. Floating point configuration

The info::fp_config enumeration tells the behavior of floating point operations on a device.

namespace sycl::info {
enum class fp_config : /* unspecified */ {
  denorm,
  inf_nan,
  round_to_nearest,
  round_to_zero,
  round_to_inf,
  fma,
  correctly_rounded_divide_sqrt,
  soft_float
};
} // namespace sycl::info

info::fp_config::denorm

Denormalized numbers are supported.


info::fp_config::inf_nan

INF and NaNs are supported.


info::fp_config::round_to_nearest

Round to nearest even rounding mode is supported.


info::fp_config::round_to_zero

Round to zero rounding mode is supported.


info::fp_config::round_to_inf

Round to positive and negative infinity rounding modes are supported.


info::fp_config::fma

IEEE754-2008 fused multiply-add is supported.


info::fp_config::correctly_rounded_divide_sqrt

Deprecated by SYCL 2020.

Divide and sqrt are correctly rounded as defined by the IEEE754 specification.


info::fp_config::soft_float

Basic floating-point operations (such as addition, subtraction, multiplication) are implemented in software.


4.6.4.7.5. Local memory type
namespace sycl::info {
enum class local_mem_type : /* unspecified */ {
  none,
  local,
  global
};
} // namespace sycl::info
4.6.4.7.6. Global memory cache type
namespace sycl::info {
enum class global_mem_cache_type : /* unspecified */ {
  none,
  read_only,
  read_write
};
} // namespace sycl::info
4.6.4.7.7. Execution capability

Deprecated by SYCL 2020.

The info::execution_capability enumeration tells the type of kernels that can be submitted to a device from the OpenCL backend.

namespace sycl::info {
enum class execution_capability : /* unspecified */ {
  exec_kernel,
  exec_native_kernel
};
} // namespace sycl::info

info::execution_capability::exec_kernel

Device can execute SYCL kernels.


info::execution_capability::exec_native_kernel

Device can execute native OpenCL kernels.


4.6.5. Queue class

The queue class encapsulates a single SYCL queue which schedules kernels on a device.

A queue can be used to submit command groups to be executed by the SYCL runtime using the queue::submit member function.

All member functions of the queue class are synchronous and errors are handled by throwing synchronous SYCL exceptions. The queue::submit member function synchronously invokes the provided command group function object (as described in Section 3.7.1.2) in the calling thread, thereby scheduling a command group for asynchronous execution. Any error in the submission of a command group is handled by throwing a synchronous SYCL exception. Any errors from the command group after it has been submitted are handled by passing asynchronous errors at specific times to an async_handler, as described in Section 4.13.

The application can wait for all command groups submitted to a queue calling queue::wait or queue::wait_and_throw.

A queue may be destroyed even when there are uncompleted commands that have been submitted to the queue. Doing so does not block. Instead, any commands that have been submitted to the queue begin execution when their requisites are satisfied, just as they would had the queue not been destroyed. Any event objects for those commands are signaled in the normal manner when the command completes. Resources associated with the queue are freed by the time the last command completes.

The queue class provides the common reference semantics as defined in Section 4.5.2.

namespace sycl {
class queue {
 public:
  explicit queue(const property_list& propList = {});

  explicit queue(const async_handler& asyncHandler,
                 const property_list& propList = {});

  template <typename DeviceSelector>
  explicit queue(const DeviceSelector& deviceSelector,
                 const property_list& propList = {});

  template <typename DeviceSelector>
  explicit queue(const DeviceSelector& deviceSelector,
                 const async_handler& asyncHandler,
                 const property_list& propList = {});

  explicit queue(const device& syclDevice, const property_list& propList = {});

  explicit queue(const device& syclDevice, const async_handler& asyncHandler,
                 const property_list& propList = {});

  template <typename DeviceSelector>
  explicit queue(const context& syclContext,
                 const DeviceSelector& deviceSelector,
                 const property_list& propList = {});

  template <typename DeviceSelector>
  explicit queue(const context& syclContext,
                 const DeviceSelector& deviceSelector,
                 const async_handler& asyncHandler,
                 const property_list& propList = {});

  explicit queue(const context& syclContext, const device& syclDevice,
                 const property_list& propList = {});

  explicit queue(const context& syclContext, const device& syclDevice,
                 const async_handler& asyncHandler,
                 const property_list& propList = {});

  /* -- common interface members -- */

  /* -- property interface members -- */

  backend get_backend() const noexcept;

  context get_context() const;

  device get_device() const;

  bool is_in_order() const;

  template <typename Param>
  typename Param::return_type get_info() const;

  template <typename Param>
  typename Param::return_type get_backend_info() const;

  template <typename T>
  event submit(T cgf);

  template <typename T>
  event submit(T cgf, const queue& secondaryQueue);

  void wait();

  void wait_and_throw();

  void throw_asynchronous();

  /* -- Shortcut functions: single_task -- */

  template <typename KernelName, typename KernelType>
  event single_task(const KernelType& kernelFunc);

  template <typename KernelName, typename KernelType>
  event single_task(event depEvent, const KernelType& kernelFunc);

  template <typename KernelName, typename KernelType>
  event single_task(const std::vector<event>& depEvents,
                    const KernelType& kernelFunc);

  /* -- Shortcut functions: parallel_for -- */

  template <typename KernelName, int Dims, typename... Rest>
  event parallel_for(range<Dims> numWorkItems, Rest&&... rest);

  template <typename KernelName, int Dims, typename... Rest>
  event parallel_for(range<Dims> numWorkItems, event depEvent, Rest&&... rest);

  template <typename KernelName, int Dims, typename... Rest>
  event parallel_for(range<Dims> numWorkItems,
                     const std::vector<event>& depEvents, Rest&&... rest);

  template <typename KernelName, int Dims, typename... Rest>
  event parallel_for(nd_range<Dims> executionRange, Rest&&... rest);

  template <typename KernelName, int Dims, typename... Rest>
  event parallel_for(nd_range<Dims> executionRange, event depEvent,
                     Rest&&... rest);

  template <typename KernelName, int Dims, typename... Rest>
  event parallel_for(nd_range<Dims> executionRange,
                     const std::vector<event>& depEvents, Rest&&... rest);

  /* -- Shortcut functions: memcpy -- */

  event memcpy(void* dest, const void* src, std::size_t numBytes);
  event memcpy(void* dest, const void* src, std::size_t numBytes, event depEvent);
  event memcpy(void* dest, const void* src, std::size_t numBytes,
               const std::vector<event>& depEvents);

  /* -- Shortcut functions: copy -- */

  template <typename T>
  event copy(const T* src, T* dest, std::size_t count);
  template <typename T>
  event copy(const T* src, T* dest, std::size_t count, event depEvent);
  template <typename T>
  event copy(const T* src, T* dest, std::size_t count,
             const std::vector<event>& depEvents);

  template <typename SrcT, int SrcDims, access_mode SrcMode, target SrcTgt,
            access::placeholder IsPlaceholder, typename DestT>
  event copy(accessor<SrcT, SrcDims, SrcMode, SrcTgt, IsPlaceholder> src,
             std::shared_ptr<DestT> dest);

  template <typename SrcT, typename DestT, int DestDims, access_mode DestMode,
            target DestTgt, access::placeholder IsPlaceholder>
  event copy(std::shared_ptr<SrcT> src,
             accessor<DestT, DestDims, DestMode, DestTgt, IsPlaceholder> dest);

  template <typename SrcT, int SrcDims, access_mode SrcMode, target SrcTgt,
            access::placeholder IsPlaceholder, typename DestT>
  event copy(accessor<SrcT, SrcDims, SrcMode, SrcTgt, IsPlaceholder> src,
             DestT* dest);

  template <typename SrcT, typename DestT, int DestDims, access_mode DestMode,
            target DestTgt, access::placeholder IsPlaceholder>
  event copy(const SrcT* src,
             accessor<DestT, DestDims, DestMode, DestTgt, IsPlaceholder> dest);

  template <typename SrcT, int SrcDims, access_mode SrcMode, target SrcTgt,
            access::placeholder IsSrcPlaceholder, typename DestT, int DestDims,
            access_mode DestMode, target DestTgt,
            access::placeholder IsDestPlaceholder>
  event copy(accessor<SrcT, SrcDims, SrcMode, SrcTgt, IsSrcPlaceholder> src,
             accessor<DestT, DestDims, DestMode, DestTgt, IsDestPlaceholder> dest);

  /* -- Shortcut functions: memset -- */

  event memset(void* ptr, int value, std::size_t numBytes);
  event memset(void* ptr, int value, std::size_t numBytes, event depEvent);
  event memset(void* ptr, int value, std::size_t numBytes,
               const std::vector<event>& depEvents);

  /* -- Shortcut functions: fill -- */

  template <typename T>
  event fill(void* ptr, const T& pattern, std::size_t count);
  template <typename T>
  event fill(void* ptr, const T& pattern, std::size_t count, event depEvent);
  template <typename T>
  event fill(void* ptr, const T& pattern, std::size_t count,
             const std::vector<event>& depEvents);

  template <typename T, int Dims, access_mode Mode, target Tgt,
            access::placeholder IsPlaceholder>
  event fill(accessor<T, Dims, Mode, Tgt, IsPlaceholder> dest, const T& src);

  /* -- Shortcut functions: prefetch -- */

  event prefetch(const void* ptr, std::size_t numBytes);
  event prefetch(const void* ptr, std::size_t numBytes, event depEvent);
  event prefetch(const void* ptr, std::size_t numBytes,
                 const std::vector<event>& depEvents);

  /* -- Shortcut functions: mem_advise -- */

  event mem_advise(const void* ptr, std::size_t numBytes, int advice);
  event mem_advise(const void* ptr, std::size_t numBytes, int advice, event depEvent);
  event mem_advise(const void* ptr, std::size_t numBytes, int advice,
                   const std::vector<event>& depEvents);

  /* -- Shortcut functions: update_host -- */

  template <typename T, int Dims, access_mode Mode, target Tgt,
            access::placeholder IsPlaceholder>
  event update_host(accessor<T, Dims, Mode, Tgt, IsPlaceholder> acc);
};
} // namespace sycl
4.6.5.1. Constructors

All queue constructors take a parameter named propList which allows the application to pass zero or more properties. These properties may specify additional effects of the constructor and may also specify exceptions that the constructor throws. See Section 4.6.5.5 for the queue properties that are defined by the core SYCL specification.


Default constructor
explicit queue(const property_list& propList = {})

Effects: Constructs a queue object using the device selected by default_selector_v. The queue’s platform is the platform that contains this device. The queue’s context is this platform’s default context as described in Section 4.6.2.


Constructor with async handler
explicit queue(const async_handler& asyncHandler,
               const property_list& propList = {})

Effects: Constructs a queue object using the device selected by default_selector_v. The queue’s platform is the platform that contains this device. The queue’s context is this platform’s default context as described in Section 4.6.2. The queue has the asynchronous error handler asyncHandler.


Constructor with device selector
template <typename DeviceSelector>
explicit queue(const DeviceSelector& deviceSelector,
               const property_list& propList = {})

Constraints: Available only when the DeviceSelector is a type that satisfies the requirements of a device selector as defined in Section 4.6.1.1.

Effects: The deviceSelector is called for every root device as described in Section 4.6.1.1, and a queue object is constructed using the device it selects. The queue’s platform is the platform that contains this device. The queue’s context is this platform’s default context as described in Section 4.6.2.


Constructor with device selector and async handler
template <typename DeviceSelector>
explicit queue(const DeviceSelector& deviceSelector,
               const async_handler& asyncHandler,
               const property_list& propList = {})

Constraints: Available only when the DeviceSelector is a type that satisfies the requirements of a device selector as defined in Section 4.6.1.1.

Effects: The deviceSelector is called for every root device as described in Section 4.6.1.1, and a queue object is constructed using the device it selects. The queue’s platform is the platform that contains this device. The queue’s context is this platform’s default context as described in Section 4.6.2. The queue has the asynchronous error handler asyncHandler.


Constructor with device
explicit queue(const device& syclDevice, const property_list& propList = {})

Effects: Constructs a queue object using the device syclDevice. The queue’s platform is the platform that contains this device. The queue’s context is this platform’s default context as described in Section 4.6.2.


Constructor with device and async handler
explicit queue(const device& syclDevice, const async_handler& asyncHandler,
               const property_list& propList = {})

Effects: Constructs a queue object using the device syclDevice. The queue’s platform is the platform that contains this device. The queue’s context is this platform’s default context as described in Section 4.6.2. The queue has the asynchronous error handler asyncHandler.


Constructor with context and device selector
template <typename DeviceSelector>
explicit queue(const context& syclContext, const DeviceSelector& deviceSelector,
               const property_list& propList = {})

Constraints: Available only when the DeviceSelector is a type that satisfies the requirements of a device selector as defined in Section 4.6.1.1.

Effects: The deviceSelector is called for every root device as described in Section 4.6.1.1, and a queue object is constructed using the device it selects. The queue’s platform is the platform that contains this device. The queue’s context is syclContext.

Throws: An exception with the errc::invalid error code if syclContext does not contain the device selected by deviceSelector.


Constructor with context, device selector, and async handler
template <typename DeviceSelector>
explicit queue(const context& syclContext, const DeviceSelector& deviceSelector,
               const async_handler& asyncHandler,
               const property_list& propList = {})

Constraints: Available only when the DeviceSelector is a type that satisfies the requirements of a device selector as defined in Section 4.6.1.1.

Effects: The deviceSelector is called for every root device as described in Section 4.6.1.1, and a queue object is constructed using the device it selects. The queue’s platform is the platform that contains this device. The queue’s context is syclContext. The queue has the asynchronous error handler asyncHandler.

Throws: An exception with the errc::invalid error code if syclContext does not contain the device selected by deviceSelector.


Constructor with context and device
explicit queue(const context& syclContext, const device& syclDevice,
               const property_list& propList = {})

Effects: Constructs a queue object using the device syclDevice. The queue’s platform is the platform that contains this device. The queue’s context is syclContext.

Throws: An exception with the errc::invalid error code unless syclDevice is contained by syclContext or is a descendent device of some device that is contained by syclContext.


Constructor with context, device, and async handler
explicit queue(const context& syclContext, const device& syclDevice,
               const async_handler& asyncHandler,
               const property_list& propList = {})

Effects: Constructs a queue object using the device syclDevice. The queue’s platform is the platform that contains this device. The queue’s context is syclContext. The queue has the asynchronous error handler asyncHandler.

Throws: An exception with the errc::invalid error code unless syclDevice is contained by syclContext or is a descendent device of some device that is contained by syclContext.


4.6.5.2. Member functions
queue::get_backend
backend get_backend() const noexcept

Returns: The SYCL backend that is associated with this queue.


queue::get_context
context get_context() const

Returns: The context that is associated with this queue.


queue::get_device
device get_device() const

Returns: The device that is associated with this queue.


queue::is_in_order
bool is_in_order() const

Returns: The same value as has_property<property::queue::in_order>().


queue::get_info
template <typename Param>
typename Param::return_type get_info() const

Constraints: Available only when Param is an information descriptor for the queue class.

Each information descriptor specifies the return value and may also specify preconditions, exceptions that are thrown, etc. See Section 4.6.5.4 for the queue information descriptors that are defined by the core SYCL specification.


queue::get_backend_info
template <typename Param>
typename Param::return_type get_backend_info() const

Constraints: Available only when Param is a backend information descriptor for the queue class.

Throws: An exception with the errc::backend_mismatch error code if the backend that corresponds with Param is different from the backend that is associated with this queue.

Each information descriptor specifies the return value and may also specify preconditions, additional exceptions that are thrown, etc.


queue::submit
template <typename T>
event submit(T cgf)

Effects: Immediately calls the command group function object cgf, which may submit no more than one command to the queue for execution on the device.

Returns: An event which represents the command which is submitted to the queue.


queue::submit (with secondary queue)
template <typename T>
event submit(T cgf, queue& secondaryQueue)

Effects: Immediately calls the command group function object cgf, which may submit no more than one command to the queue for execution on the device. On a kernel error, this command group function object may be scheduled for execution on the secondary queue secondaryQueue as described in Section 3.9.10.

Returns: An event which represents the command which is submitted to the queue. If the command is scheduled on secondaryQueue, the event is associated with that queue.


queue::wait
void wait()

Effects: Blocks the calling thread until all commands previously submitted to this queue have completed. Synchronous errors are reported through SYCL exceptions.


queue::wait_and_throw
void wait_and_throw()

Effects: Blocks the calling thread until all commands previously submitted to this queue have completed. Synchronous errors are reported through SYCL exceptions.

At least all unconsumed asynchronous errors held by this queue (or its associated context) are passed to the appropriate async_handler as described in Section 4.13.1.3.


queue::throw_asynchronous
void throw_asynchronous()

Effects: Checks to see if any unconsumed asynchronous errors have been produced by the queue and if so reports them by passing them to the async_handler associated with the queue or to the async_handler associated with the queue’s context. If no user defined asynchronous error handler is associated with the queue or its context, then an implementation-defined default async_handler is called to handle any errors, as described in Section 4.13.1.2.


4.6.5.3. Shortcut member functions

The functions described in this section are shortcuts for queue::submit that allow an application to submit a command to the queue without defining a command group function object. Each of these functions implicitly creates a command group that acts as though it calls one of the handler member functions to submit a single command. For example, queue::single_task creates a command group that acts as though it calls handler::single_task. These shortcut functions return an event object that represents the command that is submitted to the queue. In addition, some forms of the shortcut functions allow the application to pass input events, and these forms act as though the command group calls handler::depends_on with these same events.

Because there is no explicit command group function when using these shortcuts, it is not possible to create accessors for the command that is submitted. Therefore, kernels that are submitted using these shortcuts must not use accessors. Typically, applications use USM pointers instead. However, there is a special exception for non-kernel commands (e.g. shortcuts for the explicit memory copy commands). These non-kernel commands may use placeholder accessors, and the implicit command group function acts as though it calls handler::require on each of the placeholder accessors that the shortcut function uses.

The following example demonstrates the use of these shortcut functions.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
class MyKernel;

queue myQueue;
auto usmPtr = malloc_device<int>(1024, myQueue);  // USM pointer

int* data = /* pointer to some data */;
buffer buf{data, {1024}};
accessor acc{buf};  // Placeholder accessor

// Queue shortcut for a kernel invocation
myQueue.single_task<MyKernel>([=] {
  // Allowed to use USM pointers,
  // not allowed to use accessors
  usmPtr[0] = 0;
});

// Placeholder accessor will automatically be registered
myQueue.copy(data, acc);

queue::single_task
template <typename KernelName, typename KernelType>             (1)
event single_task(const KernelType& kernelFunc)

template <typename KernelName, typename KernelType>             (2)
event single_task(event depEvent, const KernelType& kernelFunc)

template <typename KernelName, typename KernelType>             (3)
event single_task(const std::vector<event>& depEvents,
                  const KernelType& kernelFunc)

Effects (1): Equivalent to calling queue::submit with a command group function that calls handler::single_task(kernelFunc).

Effects (2): Equivalent to calling queue::submit with a command group function that calls handler::depends_on(depEvent) and handler::single_task(kernelFunc).

Effects (3): Equivalent to calling queue::submit with a command group function that calls handler::depends_on(depEvents) and handler::single_task(kernelFunc).

Returns: An event which represents the command which is submitted to the queue.


queue::parallel_for
template <typename KernelName, int Dimensions, typename... Rest>        (1)
event parallel_for(range<Dimensions> numWorkItems, Rest&&... rest)

template <typename KernelName, int Dimensions, typename... Rest>        (2)
event parallel_for(range<Dimensions> numWorkItems, event depEvent,
                   Rest&&... rest)

template <typename KernelName, int Dimensions, typename... Rest>        (3)
event parallel_for(range<Dimensions> numWorkItems,
                   const std::vector<event>& depEvents, Rest&&... rest)

template <typename KernelName, int Dimensions, typename... Rest>        (4)
event parallel_for(nd_range<Dimensions> executionRange, Rest&&... rest)

template <typename KernelName, int Dimensions, typename... Rest>        (5)
event parallel_for(nd_range<Dimensions> executionRange, event depEvent,
                   Rest&&... rest)

template <typename KernelName, int Dimensions, typename... Rest>        (6)
event parallel_for(nd_range<Dimensions> executionRange,
                   const std::vector<event>& depEvents, Rest&&... rest)

Effects (1): Equivalent to calling queue::submit with a command group function that calls handler::parallel_for(numWorkItems, rest).

Effects (2): Equivalent to calling queue::submit with a command group function that calls handler::depends_on(depEvent) and handler::parallel_for(numWorkItems, rest).

Effects (3): Equivalent to calling queue::submit with a command group function that calls handler::depends_on(depEvents) and handler::parallel_for(numWorkItems, rest).

Effects (4): Equivalent to calling queue::submit with a command group function that calls handler::parallel_for(executionRange, rest).

Effects (5): Equivalent to calling queue::submit with a command group function that calls handler::depends_on(depEvent) and handler::parallel_for(executionRange, rest).

Effects (6): Equivalent to calling queue::submit with a command group function that calls handler::depends_on(depEvents) and handler::parallel_for(executionRange, rest).

Returns: An event which represents the command which is submitted to the queue.


queue::memcpy
event memcpy(void* dest, const void* src, std::size_t numBytes)                 (1)

event memcpy(void* dest, const void* src, std::size_t numBytes, event depEvent) (2)

event memcpy(void* dest, const void* src, std::size_t numBytes,                 (3)
             const std::vector<event>& depEvents)

Effects (1): Equivalent to calling queue::submit with a command group function that calls handler::memcpy(dest, src, numBytes).

Effects (2): Equivalent to calling queue::submit with a command group function that calls handler::depends_on(depEvent) and handler::memcpy(dest, src, numBytes).

Effects (3): Equivalent to calling queue::submit with a command group function that calls handler::depends_on(depEvents) and handler::memcpy(dest, src, numBytes).

Returns: An event which represents the command which is submitted to the queue.


queue::copy
template <typename T>                                                         (1)
event copy(const T* src, T* dest, std::size_t count)

template <typename T>                                                         (2)
event copy(const T* src, T* dest, std::size_t count, event depEvent)

template <typename T>                                                         (3)
event copy(const T* srct, T* dest, std::size_t count,
           const std::vector<event>& depEvents)

template <typename SrcT, int SrcDims, access_mode SrcMode, target SrcTgt,     (4)
          access::placeholder IsPlaceholder, typename DestT>
event copy(accessor<SrcT, SrcDims, SrcMode, SrcTgt, IsPlaceholder> src,
           std::shared_ptr<DestT> dest)

template <typename SrcT, typename DestT, int DestDims, access_mode DestMode,  (5)
          target DestTgt, access::placeholder IsPlaceholder>
event copy(std::shared_ptr<SrcT> src,
           accessor<DestT, DestDims, DestMode, DestTgt, IsPlaceholder> dest)

template <typename SrcT, int SrcDims, access_mode SrcMode, target SrcTgt,     (6)
          access::placeholder IsPlaceholder, typename DestT>
event copy(accessor<SrcT, SrcDims, SrcMode, SrcTgt, IsPlaceholder> src,
           DestT* dest)

template <typename SrcT, typename DestT, int DestDims, access_mode DestMode,  (7)
          target DestTgt, access::placeholder IsPlaceholder>
event copy(const SrcT* src,
           accessor<DestT, DestDims, DestMode, DestTgt, IsPlaceholder> dest)

template <typename SrcT, int SrcDims, access_mode SrcMode, target SrcTgt,     (8)
          access::placeholder IsSrcPlaceholder, typename DestT, int DestDims,
          access_mode DestMode, target DestTgt,
          access::placeholder IsDestPlaceholder>
event copy(
    accessor<SrcT, SrcDims, SrcMode, SrcTgt, IsSrcPlaceholder> src,
    accessor<DestT, DestDims, DestMode, DestTgt, IsDestPlaceholder> dest)

Effects (1): Equivalent to calling queue::submit with a command group function that calls handler::copy(src, dest, count).

Effects (2): Equivalent to calling queue::submit with a command group function that calls handler::depends_on(depEvent) and handler::copy(src, dest, count).

Effects (3): Equivalent to calling queue::submit with a command group function that calls handler::depends_on(depEvents) and handler::copy(src, dest, count).

Effects (4): Equivalent to calling queue::submit with a command group function that calls handler::require(src) and handler::copy(src, dest).

Effects (5): Equivalent to calling queue::submit with a command group function that calls handler::require(dest) and handler::copy(src, dest).

Effects (6): Equivalent to calling queue::submit with a command group function that calls handler::require(src) and handler::copy(src, dest).

Effects (7): Equivalent to calling queue::submit with a command group function that calls handler::require(dest) and handler::copy(src, dest).

Effects (8): Equivalent to calling queue::submit with a command group function that calls handler::require(src), handler::require(dest), and handler::copy(src, dest).

Returns: An event which represents the command which is submitted to the queue.


queue::memset
event memset(void* ptr, int value, std::size_t numBytes)                 (1)

event memset(void* ptr, int value, std::size_t numBytes, event depEvent) (2)

event memset(void* ptr, int value, std::size_t numBytes,                 (3)
             const std::vector<event>& depEvents)

Effects (1): Equivalent to calling queue::submit with a command group function that calls handler::memset(ptr, value, numBytes).

Effects (2): Equivalent to calling queue::submit with a command group function that calls handler::depends_on(depEvent) and handler::memset(ptr, value, numBytes).

Effects (3): Equivalent to calling queue::submit with a command group function that calls handler::depends_on(depEvents) and handler::memcpy(ptr, value, numBytes).

Returns: An event which represents the command which is submitted to the queue.


queue::fill
template <typename T>                                                      (1)
event fill(void* ptr, const T& pattern, std::size_t count)

template <typename T>                                                      (2)
event fill(void* ptr, const T& pattern, std::size_t count, event depEvent)

template <typename T>                                                      (3)
event fill(void* ptr, const T& pattern, std::size_t count,
           const std::vector<event>& depEvents)

template <typename T, int Dims, access_mode Mode, target Tgt,              (4)
          access::placeholder IsPlaceholder>
event fill(accessor<T, Dims, Mode, Tgt, IsPlaceholder> dest, const T& src)

Effects (1): Equivalent to calling queue::submit with a command group function that calls handler::fill(ptr, pattern, count).

Effects (2): Equivalent to calling queue::submit with a command group function that calls handler::depends_on(depEvent) and handler::fill(ptr, pattern, count).

Effects (3): Equivalent to calling queue::submit with a command group function that calls handler::depends_on(depEvents) and handler::fill(ptr, pattern, count).

Effects (4): Equivalent to calling queue::submit with a command group function that calls handler::require(dest) and handler::fill(dest, src).

Returns: An event which represents the command which is submitted to the queue.


queue::prefetch
event prefetch(const void* ptr, std::size_t numBytes)                                      (1)

event prefetch(const void* ptr, std::size_t numBytes, event depEvent)                      (2)

event prefetch(const void* ptr, std::size_t numBytes, const std::vector<event>& depEvents) (3)

Effects (1): Equivalent to calling queue::submit with a command group function that calls handler::prefetch(ptr, numBytes).

Effects (2): Equivalent to calling queue::submit with a command group function that calls handler::depends_on(depEvent) and handler::prefetch(ptr, numBytes).

Effects (3): Equivalent to calling queue::submit with a command group function that calls handler::depends_on(depEvents) and handler::prefetch(ptr, numBytes).

Returns: An event which represents the command which is submitted to the queue.


queue::mem_advise
event mem_advise(const void* ptr, std::size_t numBytes, int advice)                 (1)

event mem_advise(const void* ptr, std::size_t numBytes, int advice, event depEvent) (2)

event mem_advise(const void* ptr, std::size_t numBytes, int advice,                 (3)
                 const std::vector<event>& depEvents)

Effects (1): Equivalent to calling queue::submit with a command group function that calls handler::mem_advise(ptr, numBytes, advice).

Effects (2): Equivalent to calling queue::submit with a command group function that calls handler::depends_on(depEvent) and handler::mem_advise(ptr, numBytes, advice).

Effects (3): Equivalent to calling queue::submit with a command group function that calls handler::depends_on(depEvents) and handler::mem_advise(ptr, numBytes, advice).

Returns: An event which represents the command which is submitted to the queue.


queue::update_host
template <typename T, int Dims, access_mode Mode, target Tgt,
          access::placeholder IsPlaceholder>
event update_host(accessor<T, Dims, Mode, Tgt, IsPlaceholder> acc)

Effects: Equivalent to calling queue::submit with a command group function that calls handler::require(acc) and handler::update_host(acc).

Returns: An event which represents the command which is submitted to the queue.


4.6.5.4. Information descriptors

This section describes the information descriptors that can be used as the Param template parameter to queue::get_info. When the description has a Returns, Throws, etc. paragraph, this indicates the value returned by or the exceptions thrown by the queue::get_info function.


info::queue::context
namespace sycl::info::queue {
struct context {
  using return_type = context;
};
} // namespace sycl::info::queue

Remarks: Template parameter to queue::get_info.

Returns: The context that is associated with this queue.


info::queue::device
namespace sycl::info::queue {
struct device {
  using return_type = device;
};
} // namespace sycl::info::queue

Remarks: Template parameter to queue::get_info.

Returns: The device that is associated with this queue.


4.6.5.5. Properties

This section describes the properties that can be passed in the propList parameter of the queue constructors.


property::queue::enable_profiling
namespace sycl::property::queue {
struct enable_profiling {
  enable_profiling();  (1)
};
} // namespace sycl::property::queue

When a queue is constructed with this property, the implementation captures profiling information for the command groups that are submitted to this queue. Applications can retrieve this profiling information by calling event::get_profiling_info on the event that is returned when submitting the command group. If the queue’s associated device does not have aspect::queue_profiling, passing this property to the queue’s constructor causes the constructor to throw a synchronous exception with the errc::feature_not_supported error code.

Effects (1): Constructs an enable_profiling property object.


property::queue::in_order
namespace sycl::property::queue {
struct in_order {
  in_order();  (1)
};
} // namespace sycl::property::queue

When a queue is constructed with this property, commands that are submitted to the queue are guaranteed to execute in the order in which they are submitted, as if there is an implicit dependency on the previous command that was submitted to the same queue. The in_order property does not provide any guarantee about the order of commands submitted to other queues with respect to commands submitted to this queue.

Effects (1): Constructs an in_order property object.


4.6.5.6. Error handling

Queue errors come in two forms:

  • Synchronous errors are those that we would expect to be reported directly at the point of waiting on an event, and hence waiting for a queue to complete, as well as any immediate errors reported by enqueuing work onto a queue. Such errors are reported through C++ exceptions.

  • Asynchronous errors are those that are produced or detected after associated host API calls have returned (so can’t be thrown as exceptions by the API call), and that are handled by an async_handler through which the errors are reported. Handling of asynchronous errors from a queue occurs at specific times, as described by Section 4.13.

Note that if there are asynchronous errors to be processed when a queue is destroyed, the handler is called and this might delay or block the destruction, according to the behavior of the handler.

4.6.6. Event class

An event in SYCL is an object that represents the status of an operation that is being executed by the SYCL runtime.

Typically in SYCL, data dependency and execution order is handled implicitly by the SYCL runtime. However, in some circumstances developers want fine grain control of the execution, or want to retrieve properties of a command that is running.

Note that, although an event represents the status of a particular operation, the dependencies of a certain event can be used to keep track of multiple steps required to block on the results of said operation.

A SYCL event is returned by the submission of a command group. The dependencies of the event returned via the submission of the command group are the implementation-defined commands associated with the command group execution.

The SYCL event class provides the common reference semantics (see Section 4.5.2).

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
namespace sycl {

class event {
 public:
  event();

  /* -- common interface members -- */

  backend get_backend() const noexcept;

  std::vector<event> get_wait_list();

  void wait();

  static void wait(const std::vector<event>& eventList);

  void wait_and_throw();

  static void wait_and_throw(const std::vector<event>& eventList);

  template <typename Param> typename Param::return_type get_info() const;

  template <typename Param>
  typename Param::return_type get_backend_info() const;

  template <typename Param>
  typename Param::return_type get_profiling_info() const;
};

} // namespace sycl
4.6.6.1. Constructors
Default constructor
event()

Effects: Constructs an event that is immediately ready. The event has no dependencies and no associated commands. Waiting on this event will return immediately and querying its status will return info::event_command_status::complete.

Remarks: The event is constructed as though it were created from a default-constructed queue. Therefore, its backend is the same as the backend of the device selected by default_selector_v.


4.6.6.2. Member functions
event::get_backend
backend get_backend() const noexcept

Returns: The SYCL backend associated with this event.


event::get_wait_list
std::vector<event> get_wait_list()

Returns: The list of events that this event waits for in the dependence graph. Only direct dependencies are returned, and not transitive dependencies that direct dependencies wait on.

Remarks: Whether already completed events are included in the returned list is implementation-defined. If the event is associated with a command submitted to an in-order queue, then it is implementation-defined whether this function returns any dependent events or if it returns an empty vector.


event::wait
void wait()

Effects: Blocks until all commands associated with this event and any dependent events have completed.


event::wait_and_throw
void wait_and_throw()

Effects:

  • Blocks until all commands associated with this event and any dependent events have completed.

  • At least all unconsumed asynchronous errors held by queues (or their associated contexts) which were used to enqueue commands associated with this event and any dependent events are passed to the appropriate async_handler as described in Section 4.13.1.3.

[Note: This behavior is equivalent to calling queue::throw_asynchronous on the queue associated with this event and any dependent events. — end note]


event::get_info
template <typename Param>
typename Param::return_type get_info() const

Constraints: Available only when Param is an information descriptor for the event class.

Each information descriptor specifies the return value and may also specify preconditions, exceptions that are thrown, etc. See Section 4.6.6.4 for the event information descriptors that are defined by the core SYCL specification.


event::get_backend_info
template <typename Param>
typename Param::return_type get_backend_info() const

Constraints: Available only when Param is a backend information descriptor for the event class.

Throws: An exception with the errc::backend_mismatch error code if the backend that corresponds with Param is different from the backend that is associated with this event.

Each information descriptor specifies the return value and may also specify preconditions, additional exceptions that are thrown, etc.


event::get_profiling_info
template <typename Param>
typename Param::return_type get_profiling_info() const

Constraints: Available only when Param is a profiling information descriptor for the event class.

Effects: If the requested profiling information is unavailable when get_profiling_info is called due to incompletion of command groups associated with the event, then the call to get_profiling_info will block until the requested profiling information is available.

[Note: An example is asking for info::event_profiling::command_end when the associated command group action has yet to finish execution. — end note]

Throws: An exception with the errc::invalid error code if the SYCL queue that submitted the command group that this event is associated with was not constructed with the property::queue::enable_profiling property.

Each profiling information descriptor specifies the return value and may also specify preconditions, additional exceptions that are thrown, etc. See Section 4.6.6.5 for the profiling information descriptors that are defined by the core SYCL specification.


4.6.6.3. Static member functions
event::wait (with event list)
static void wait(const std::vector<event>& eventList)

Effects: Behaves as if calling event::wait on each event in eventList.


event::wait_and_throw (with event list)
static void wait_and_throw(const std::vector<event>& eventList)

Effects: Behaves as if calling event::wait_and_throw on each event in eventList.


4.6.6.4. Information descriptors

This section describes the information descriptors that can be used as the Param template parameter to event::get_info. When the description has a Returns, Throws, etc. paragraph, this indicates the value returned by or the exceptions thrown by the event::get_info function.


info::event::command_execution_status
namespace sycl::info::event {
struct command_execution_status {
  using return_type = info::event_command_status;
};
} // namespace sycl::info::event

Returns: The event status of the command group and contained action (e.g. kernel invocation) associated with this event. The value returned is one of the following:

Remarks: Template parameter to event::get_info.


4.6.6.5. Profiling information descriptors

This section describes the profiling information descriptors that can be used as the Param template parameter to event::get_profiling_info.

Each profiling descriptor returns a 64-bit timestamp that represents the number of nanoseconds that have elapsed since some implementation-defined timebase. All events that share the same backend are guaranteed to share the same timebase, and therefore the difference between two timestamps from the same backend yields the number of nanoseconds that have elapsed between those events.

When the description has a Returns, Throws, etc. paragraph, this indicates the value returned by or the exceptions thrown by the event::get_profiling_info function.


info::event_profiling::command_submit
namespace sycl::info::event_profiling {
struct command_submit {
  using return_type = std::uint64_t;
};
} // namespace sycl::info::event_profiling

Returns: The timestamp corresponding to when the associated command group was submitted to the queue.

Remarks:


info::event_profiling::command_start
namespace sycl::info::event_profiling {
struct command_start {
  using return_type = std::uint64_t;
};
} // namespace sycl::info::event_profiling

Effects: Querying this profiling descriptor blocks until the event’s state becomes either info::event_command_status::running or info::event_command_status::complete.

Returns: The timestamp corresponding to when the action associated with the command group (e.g., kernel invocation) started executing on the device.

Remarks:

[Note: Implementations are encouraged to return a timestamp that is as close as possible to the point when the action starts running on the device, but there is no specific accuracy that is guaranteed. — end note]


info::event_profiling::command_end
namespace sycl::info::event_profiling {
struct command_end {
  using return_type = std::uint64_t;
};
} // namespace sycl::info::event_profiling

Effects: Querying this profiling descriptor blocks until the event’s state becomes info::event_command_status::complete.

Returns: The timestamp corresponding to when the action associated with the command group (e.g., kernel invocation) finished executing on the device.

Remarks:


4.6.6.6. Other enumerations
4.6.6.6.1. Event command status
namespace sycl::info {
enum class event_command_status : /* unspecified */ {
  submitted,
  running,
  complete
};
} // namespace sycl::info
info::event_command_status::submitted

Indicates that the command has been submitted to the SYCL queue but has not yet started running on the device.


info::event_command_status::running

Indicates that the command has started running on the device but has not yet completed.


info::event_command_status::complete

Indicates that the command has finished running on the device. Attempting to wait on such an event will not block.

Synchronization: When event::get_info returns this value, the synchronization is equivalent to event::wait.

4.7. Data access and storage in SYCL

In SYCL, when using buffers and images, data storage and access are handled by separate classes. Buffers and images handle storage and ownership of the data, whereas accessors handle access to the data. Buffers and images in SYCL can be bound to more than one device or context, including across different SYCL backends. They also handle ownership of the data, while allowing exception handling for blocking and non-blocking data transfers. Accessors manage data transfers between the host and all of the devices in the system, as well as tracking of data dependencies.

Zero-sized buffers and accessors are permitted, but attempting to access data within them produces undefined behavior, similar to dereferencing a null pointer in C++. Note that zero-sized accessors can be created in several ways: by creating an accessor from a zero-sized buffer, by creating an accessor with a zero-sized buffer sub-range, or by creating an accessor with its default constructor.

When using USM allocations, data storage is managed by USM allocation functions, and data access is via pointers. See Section 4.8 for greater detail.

4.7.1. Host allocation

A SYCL runtime may need to allocate temporary objects on the host to handle some operations (such as copying data from one context to another). Allocation on the host is managed using an allocator object, following the standard C++ allocator class definition. The default allocator for memory objects is implementation-defined, but the user can supply their own allocator class.

1
2
3
{
    buffer<int, 1, UserDefinedAllocator<int>> b(d);
}

Note that if the runtime requires host memory (e.g., when moving data across SYCL backend contexts), but the allocator fails to allocate the memory, then the runtime will raise an error.

In some cases, the implementation may retain a copy of the allocator object even after the buffer is destroyed. For example, this can happen when the buffer object is destroyed before commands using accessors to the buffer have completed. Therefore, the application must be prepared for calls to the allocator even after the buffer is destroyed.

If the application needs to know when the implementation has destroyed all copies of the allocator, it can maintain a reference count within the allocator.

The definition of allocators extends the current functionality of SYCL, ensuring that users can define allocator functions for specific hardware or certain complex shared memory mechanisms (e.g. NUMA), and improves interoperability with STL-based libraries (e.g, Intel’s TBB provides an allocator).

4.7.1.1. Default allocators

A default allocator is always defined by the implementation. The default allocator for const buffers will remove the const-ness of the type (therefore, the default allocator for a buffer of type const int will be an Allocator<int>). This implies that host accessors will not share memory with the pointer given by the user in the buffer/image constructor, but will use the memory returned by the Allocator itself for that purpose. The user can implement an allocator that returns the same address as the one passed in the buffer constructor, but it is the responsibility of the user to handle the potential race conditions.

Table 15. SYCL Default Allocators
Allocators Description
template <class T> buffer_allocator

It is the default buffer allocator used by the runtime, when no allocator is defined by the user. Meets the C++ named requirement Allocator. A buffer of data type const T uses buffer_allocator<T> by default.

image_allocator

It is the default allocator used by the runtime for the SYCL unsampled_image and sampled_image classes when no allocator is provided by the user. The image_allocator is required to allocate in elements of std::byte.

See Section 4.7.5 for details of using manual synchronization to avoid data races between host and device.

4.7.2. Buffers

The buffer class defines a shared array of one, two or three dimensions that can be used by the SYCL kernel and has to be accessed using accessor classes. Buffers are templated on both the type of their data, and the number of dimensions that the data is stored and accessed through.

A buffer does not map to only one underlying backend object, and all SYCL backend memory objects may be temporary for use within a command group on a specific device.

The underlying data type of a buffer T must be device copyable as defined in Section 3.13.1. Some overloads of the buffer constructor initialize the buffer contents by copying objects from host memory while other overloads construct the buffer without copying objects from the host. For the overloads that do not copy host objects, the initial state of the objects in the buffer depends on whether T is an implicit-lifetime type (as defined in the C++ core language). If T is an implicit-lifetime type, objects of that type are implicitly created in the buffer with indeterminate values. For other types, these constructor overloads merely allocate uninitialized memory, and the application is responsible for constructing objects by calling placement-new and for destroying them later by manually calling the object’s destructor.

For the overloads that do copy objects from host memory, the hostData pointer must point to at least N bytes of memory where N is sizeof(T) * bufferRange.size(). If N is zero, hostData is permitted to be a null pointer.

A SYCL buffer can construct an instance of a SYCL buffer that reinterprets the original SYCL buffer with a different type, dimensionality and range using the member function reinterpret. The reinterpreted SYCL buffer that is constructed must behave as though it were a copy of the SYCL buffer that constructed it (see Section 4.5.2) with the exception that the type, dimensionality and range of the reinterpreted SYCL buffer must reflect the type, dimensionality and range specified when calling the reinterpret member function. By extension of this, the class member types value_type, reference and const_reference, and the member functions get_range() and size() of the reinterpreted SYCL buffer must reflect the new type, dimensionality and range. The data that the original SYCL buffer and the reinterpreted SYCL buffer manage remains unaffected, though the representation of the data when accessed through the reinterpreted SYCL buffer may alter to reflect the new type, dimensionality and range. It is important to note that a reinterpreted SYCL buffer is a copy of the original SYCL buffer only, and not a new SYCL buffer. Constructing more than one SYCL buffer managing the same host pointer is still undefined behavior.

The SYCL buffer class template provides the common reference semantics (see Section 4.5.2).

The SYCL buffer class template takes a template parameter AllocatorT for specifying an allocator which is used by the SYCL runtime when allocating temporary memory on the host. If no template argument is provided, then the default allocator for the SYCL buffer class buffer_allocator<T> will be used (see Section 4.7.1.1).

namespace sycl {
template <typename T, int Dimensions = 1,
          typename AllocatorT = buffer_allocator<std::remove_const_t<T>>>
class buffer {
 public:
  using value_type = T;
  using reference = value_type&;
  using const_reference = const value_type&;
  using allocator_type = AllocatorT;

  buffer(const range<Dimensions>& bufferRange, AllocatorT allocator,
         const property_list& propList = {});

  buffer(const range<Dimensions>& bufferRange,
         const property_list& propList = {});

  buffer(T* hostData, const range<Dimensions>& bufferRange,
         AllocatorT allocator, const property_list& propList = {});

  buffer(T* hostData, const range<Dimensions>& bufferRange,
         const property_list& propList = {});

  buffer(const T* hostData, const range<Dimensions>& bufferRange,
         AllocatorT allocator, const property_list& propList = {});

  buffer(const T* hostData, const range<Dimensions>& bufferRange,
         const property_list& propList = {});

  /* Available only if Container is a contiguous container:
       - std::data(container) and std::size(container) are well formed
       - return type of std::data(container) is convertible to T*
     and Dimensions == 1 */
  template <typename Container>
  buffer(Container& container, AllocatorT allocator,
         const property_list& propList = {});

  /* Available only if Container is a contiguous container:
       - std::data(container) and std::size(container) are well formed
       - return type of std::data(container) is convertible to T*
     and Dimensions == 1 */
  template <typename Container>
  buffer(Container& container, const property_list& propList = {});

  buffer(const std::shared_ptr<T>& hostData,
         const range<Dimensions>& bufferRange, AllocatorT allocator,
         const property_list& propList = {});

  buffer(const std::shared_ptr<T>& hostData,
         const range<Dimensions>& bufferRange,
         const property_list& propList = {});

  buffer(const std::shared_ptr<T[]>& hostData,
         const range<Dimensions>& bufferRange, AllocatorT allocator,
         const property_list& propList = {});

  buffer(const std::shared_ptr<T[]>& hostData,
         const range<Dimensions>& bufferRange,
         const property_list& propList = {});

  template <typename InputIterator>
  buffer(InputIterator first, InputIterator last, AllocatorT allocator,
         const property_list& propList = {});

  template <typename InputIterator>
  buffer(InputIterator first, InputIterator last,
         const property_list& propList = {});

  buffer(buffer& b, const id<Dimensions>& baseIndex,
         const range<Dimensions>& subRange);

  /* -- common interface members -- */

  /* -- property interface members -- */

  range<Dimensions> get_range() const;

  std::size_t byte_size() const noexcept;

  std::size_t size() const noexcept;

  // Deprecated
  std::size_t get_count() const;

  // Deprecated
  std::size_t get_size() const;

  AllocatorT get_allocator() const;

  template <access_mode Mode = access_mode::read_write,
            target Targ = target::device>
  accessor<T, Dimensions, Mode, Targ> get_access(handler& commandGroupHandler);

  // Deprecated
  template <access_mode Mode>
  accessor<T, Dimensions, Mode, target::host_buffer> get_access();

  template <access_mode Mode = access_mode::read_write,
            target Targ = target::device>
  accessor<T, Dimensions, Mode, Targ>
  get_access(handler& commandGroupHandler, range<Dimensions> accessRange,
             id<Dimensions> accessOffset = {});

  // Deprecated
  template <access_mode Mode>
  accessor<T, Dimensions, Mode, target::host_buffer>
  get_access(range<Dimensions> accessRange, id<Dimensions> accessOffset = {});

  template <typename... Ts> auto get_access(Ts...);

  template <typename... Ts> auto get_host_access(Ts...);

  template <typename Destination = std::nullptr_t>
  void set_final_data(Destination finalData = nullptr);

  void set_write_back(bool flag = true);

  bool is_sub_buffer() const;

  template <typename ReinterpretT, int ReinterpretDim>
  buffer<ReinterpretT, ReinterpretDim,
         typename std::allocator_traits<AllocatorT>::template rebind_alloc<
             ReinterpretT>>
  reinterpret(range<ReinterpretDim> reinterpretRange) const;

  // Only available when ReinterpretDim == 1
  // or when (ReinterpretDim == Dimensions) &&
  //         (sizeof(ReinterpretT) == sizeof(T))
  template <typename ReinterpretT, int ReinterpretDim = Dimensions>
  buffer<ReinterpretT, ReinterpretDim,
         typename std::allocator_traits<AllocatorT>::template rebind_alloc<
             ReinterpretT>>
  reinterpret() const;
};

// Deduction guides
template <typename InputIterator, typename AllocatorT>
buffer(InputIterator, InputIterator, AllocatorT, const property_list& = {})
    -> buffer<typename std::iterator_traits<InputIterator>::value_type, 1,
              AllocatorT>;

template <typename InputIterator>
buffer(InputIterator, InputIterator, const property_list& = {})
    -> buffer<typename std::iterator_traits<InputIterator>::value_type, 1>;

template <typename T, int Dimensions, typename AllocatorT>
buffer(const T*, const range<Dimensions>&, AllocatorT,
       const property_list& = {}) -> buffer<T, Dimensions, AllocatorT>;

template <typename T, int Dimensions>
buffer(const T*, const range<Dimensions>&, const property_list& = {})
    -> buffer<T, Dimensions>;

template <typename Container, typename AllocatorT>
buffer(Container&, AllocatorT, const property_list& = {})
    -> buffer<typename Container::value_type, 1, AllocatorT>;

template <typename Container>
buffer(Container&, const property_list& = {})
    -> buffer<typename Container::value_type, 1>;

} // namespace sycl
4.7.2.1. Constructors

All buffer constructors take a parameter named propList which allows the application to pass zero or more properties. These properties may specify additional effects of the constructor and resulting buffer object. See Section 4.7.2.3 for the buffer properties that are defined by the core SYCL specification.

Construct with uninitialized memory
buffer(const range<Dimensions>& bufferRange, AllocatorT allocator,  (1)
       const property_list& propList = {});

buffer(const range<Dimensions>& bufferRange,                        (2)
       const property_list& propList = {});

Effects (1): Construct a SYCL buffer instance with uninitialized memory. The constructed SYCL buffer will use the allocator parameter provided when allocating memory on the host. The range of the constructed SYCL buffer is specified by the bufferRange parameter provided. Data is not written back to the host on destruction of the buffer unless the buffer has a valid non-null pointer specified via the member function set_final_data(). Zero or more properties can be provided to the constructed SYCL buffer via an instance of property_list.

Effects (2): Equivalent to buffer(bufferRange, AllocatorT{}, propList).


Construct with host memory
buffer(T* hostData, const range<Dimensions>& bufferRange,          (1)
       AllocatorT allocator, const property_list& propList = {});

buffer(T* hostData, const range<Dimensions>& bufferRange,          (2)
       const property_list& propList = {});

buffer(const T* hostData, const range<Dimensions>& bufferRange,    (3)
       AllocatorT allocator, const property_list& propList = {});

buffer(const T* hostData, const range<Dimensions>& bufferRange,    (4)
       const property_list& propList = {});

Effects (1): Construct a SYCL buffer instance with the hostData parameter provided. The buffer is initialized with the memory specified by hostData, and the buffer assumes exclusive access to this memory for the duration of its lifetime. The constructed SYCL buffer will use the allocator parameter provided when allocating memory on the host. The range of the constructed SYCL buffer is specified by the bufferRange parameter provided. Zero or more properties can be provided to the constructed SYCL buffer via an instance of property_list.

Effects (2): Equivalent to buffer(hostData, bufferRange, AllocatorT{}, propList).

Effects (3): Construct a SYCL buffer instance with the hostData parameter provided. The buffer assumes exclusive access to this memory for the duration of its lifetime. The constructed SYCL buffer will use the allocator parameter provided when allocating memory on the host. The range of the constructed SYCL buffer is specified by the bufferRange parameter provided. Zero or more properties can be provided to the constructed SYCL buffer via an instance of property_list.

Effects (4): Equivalent to buffer(hostData, bufferRange, AllocatorT{}, propList).

Remarks: When hostData is a pointer to a const-qualified type, the buffer will not write back to any host memory unless requested via the member function set_final_data().


Construct from a container
template <typename Container>                                      (1)
buffer(Container& container, AllocatorT allocator,
       const property_list& propList = {});

template <typename Container>                                      (2)
buffer(Container& container, const property_list& propList = {});

Preconditions: container is a contiguous container.

Constraints: Available only when:

  • std::data(container) and std::size(container) are well formed;

  • the return type of std::data(container) is convertible to T*; and

  • Dimensions == 1.

Effects (1): Construct a one dimensional SYCL buffer instance from the elements starting at std::data(container) and containing std::size(container) number of elements. The buffer is initialized with the contents of container, and the buffer assumes exclusive access to container for the duration of its lifetime. The constructed SYCL buffer will use the allocator parameter provided when allocating memory on the host. Zero or more properties can be provided to the constructed SYCL buffer via an instance of property_list.

Effects (2): Equivalent to buffer(container, AllocatorT{}, propList).

Remarks: Data is written back to container before the completion of buffer destruction if the return type of std::data(container) is not const.


Construct from memory owned by a shared pointer
  buffer(const std::shared_ptr<T>& hostData,                          (1)
         const range<Dimensions>& bufferRange, AllocatorT allocator,
         const property_list& propList = {});

  buffer(const std::shared_ptr<T>& hostData,                          (2)
         const range<Dimensions>& bufferRange,
         const property_list& propList = {});

  buffer(const std::shared_ptr<T[]>& hostData,                        (3)
         const range<Dimensions>& bufferRange, AllocatorT allocator,
         const property_list& propList = {});

  buffer(const std::shared_ptr<T[]>& hostData,                        (4)
         const range<Dimensions>& bufferRange,
         const property_list& propList = {});

Effects (1): When hostData is not empty, construct a SYCL buffer with the contents of its stored pointer. The buffer assumes exclusive access to this memory for the duration of its lifetime. The buffer also creates its own internal copy of the shared_ptr that shares ownership of the hostData memory, which means the application can safely release ownership of this shared_ptr when the constructor returns. When hostData is empty, construct a SYCL buffer with uninitialized memory. The constructed SYCL buffer will use the allocator parameter provided when allocating memory on the host. The range of the constructed SYCL buffer is specified by the bufferRange parameter provided. Zero or more properties can be provided to the constructed SYCL buffer via an instance of property_list.

Effects (2): Equivalent to buffer(hostData, bufferRange, AllocatorT{}, propList).

Effects (3): When hostData is not empty, construct a SYCL buffer with the contents of its stored pointer. The buffer assumes exclusive access to this memory for the duration of its lifetime. The buffer also creates its own internal copy of the shared_ptr that shares ownership of the hostData memory, which means the application can safely release ownership of this shared_ptr when the constructor returns. When hostData is empty, construct a SYCL buffer with uninitialized memory. The constructed SYCL buffer will use the allocator parameter provided when allocating memory on the host. The range of the constructed SYCL buffer is specified by the bufferRange parameter provided. Zero or more properties can be provided to the constructed SYCL buffer via an instance of property_list.

Effects (4): Equivalent to buffer(hostData, bufferRange, AllocatorT{}, propList).


Construct from iterators
template <typename InputIterator>                                      (1)
buffer(InputIterator first, InputIterator last, AllocatorT allocator,
       const property_list& propList = {});

template <typename InputIterator>                                      (2)
buffer(InputIterator first, InputIterator last,
       const property_list& propList = {});

Effects (1): Create a new allocated 1D buffer initialized from the given elements ranging from first up to one before last. The data is copied to an intermediate memory position by the runtime. Data is not written back to the same iterator set provided. However, if the buffer has a valid non-const iterator specified via the member function set_final_data(), data will be copied back to that iterator. The constructed SYCL buffer will use the allocator parameter provided when allocating memory on the host. Zero or more properties can be provided to the constructed SYCL buffer via an instance of property_list.

Effects (2): Equivalent to buffer(first, last, AllocatorT{}, propList).

Remarks: The buffer will not write back to any host memory unless requested via the member function set_final_data().


Construct sub-buffer
buffer(buffer& b, const id<Dimensions>& baseIndex,
       const range<Dimensions>& subRange);

Effects: Create a new sub-buffer without allocation to have separate accessors later. b is the buffer with the real data. baseIndex specifies the origin of the sub-buffer inside the buffer b. subRange specifies the size of the sub-buffer.

The origin (based on baseIndex) of the sub-buffer being constructed must be a multiple of the memory base address alignment of each SYCL device which accesses data from the buffer. This value is retrievable via the SYCL device class info query info::device::mem_base_addr_align. Violating this requirement causes the implementation to throw an exception with the errc::invalid error code from the accessor constructor (if the accessor is not a placeholder) or from handler::require() (if the accessor is a placeholder). If the accessor is bound to a command group with a secondary queue, the sub-buffer’s alignment must be compatible with both the primary queue’s device and the secondary queue’s device. If the implementation supports secondary queue fallback, it also throws this exception if the sub-buffer’s alignment is not compatible with the secondary queue’s device.

Throws:

  • An exception with the errc::invalid error code if b is a sub-buffer.

  • An exception with the errc::invalid error code if the subrange formed by baseIndex and subRange is not a contiguous region of b.

  • An exception with the errc::invalid error code if the sum of baseIndex and subRange in any dimension exceeds the parent buffer b size (bufferRange) in that dimension.


4.7.2.2. Member functions
buffer::get_range
range<Dimensions> get_range() const;

Returns: A range representing the number of elements in the buffer.


buffer::byte_size
std::size_t byte_size() const noexcept;

Returns: The size of the buffer storage in bytes. Equal to size()*sizeof(T).


buffer::size
std::size_t size() const noexcept;

Returns: The total number of elements in the buffer. Equal to get_range()[0] * ... * get_range()[Dimensions-1].


buffer::get_count
std::size_t get_count() const;

Deprecated by SYCL 2020.

Effects: Equivalent to return size().


buffer::get_size
std::size_t get_size() const;

Deprecated by SYCL 2020.

Effects: Equivalent to return byte_size().


buffer::get_allocator
AllocatorT get_allocator() const;

Returns: The allocator provided to the buffer.


buffer::get_access
template <access_mode Mode = access_mode::read_write,                          (1)
          target Targ = target::device>
accessor<T, Dimensions, Mode, Targ> get_access(handler& commandGroupHandler);

template <access_mode Mode = access_mode::read_write,                          (2)
          target Targ = target::device>
accessor<T, Dimensions, Mode, Targ>
get_access(handler& commandGroupHandler, range<Dimensions> accessRange,
           id<Dimensions> accessOffset = {});

template <typename... Ts> auto get_access(Ts...);                              (3)

template <typename... Ts> auto get_host_access(Ts...);                         (4)

Constraints (1)-(2): Available only when Targ is target::device, target::constant_buffer or target::host_task.

Returns (1): A valid accessor to the buffer with the specified access mode and target in the command group buffer.

Returns (2): A valid accessor to the buffer with the specified access mode and target in the command group buffer. The accessor is a ranged accessor, where the range starts at the given offset from the beginning of the buffer.

Returns (3): A valid accessor as if constructed via passing the buffer and all provided arguments to the accessor constructor.

Returns (4): A valid host_accessor as if constructed via passing the buffer and all provided arguments to the host_accessor constructor.

Throws (2): An exception with the errc::invalid error code if the sum of accessRange and accessOffset exceeds the range of the buffer in any dimension.


Deprecated buffer::get_access
template <access_mode Mode>                                                   (1)
accessor<T, Dimensions, Mode, target::host_buffer> get_access();

template <access_mode Mode>                                                   (2)
accessor<T, Dimensions, Mode, target::host_buffer>
get_access(range<Dimensions> accessRange, id<Dimensions> accessOffset = {});

Deprecated in SYCL 2020. Use get_host_access() instead.

Returns (1): A valid host accessor to the buffer with the specified access mode.

Returns (2): A valid host accessor to the buffer with the specified access mode. The accessor is a ranged accessor, where the range starts at the given offset from the beginning of the buffer.

Throws (2): An exception with the errc::invalid error code if the sum of accessRange and accessOffset exceeds the range of the buffer in any dimension.


buffer::set_final_data
template <typename Destination = std::nullptr_t>
void set_final_data(Destination finalData = nullptr);

Effects: The finalData points to where the outcome of all the buffer processing is going to be copied to at destruction time, if the buffer was involved with a write accessor. Destination can be either an output iterator or a std::weak_ptr<T>. Note that a raw pointer is a special case of output iterator and thus defines the host memory to which the result is to be copied. In the case of a weak pointer, the output is not updated if the weak pointer has expired. If Destination is std::nullptr_t, then the copy back will not happen.


buffer::set_write_back
void set_write_back(bool flag = true);

Effects: Dynamically forces or cancels the write-back of the data of a buffer on destruction according to the value of flag. Forcing the write-back is similar to what happens during a normal write-back as described in Section 4.7.2.4 and Section 4.7.4. If there is nowhere to write-back, using this function does not have any effect.


buffer::is_sub_buffer
bool is_sub_buffer() const;

Returns: true if this SYCL buffer is a sub-buffer and false otherwise.


buffer::reinterpret
template <typename ReinterpretT, int ReinterpretDim>                       (1)
buffer<ReinterpretT, ReinterpretDim,
       typename std::allocator_traits<AllocatorT>::template rebind_alloc<
           ReinterpretT>>
reinterpret(range<ReinterpretDim> reinterpretRange) const;

template <typename ReinterpretT, int ReinterpretDim = Dimensions>          (2)
buffer<ReinterpretT, ReinterpretDim,
       typename std::allocator_traits<AllocatorT>::template rebind_alloc<
           ReinterpretT>>
reinterpret() const;

Constraints (2): Available when (ReinterpretDim == 1) or when ((ReinterpretDim == Dimensions) && (sizeof(ReinterpretT) == sizeof(T))).

Effects: Creates a reinterpreted SYCL buffer. The buffer object being reinterpreted can be a SYCL sub-buffer that was created from a SYCL buffer. Reinterpreting a sub-buffer provides a reinterpreted view of the sub-buffer only, and does not change the offset or size of the sub-buffer view (in bytes) relative to the parent buffer.

Returns (1): A reinterpreted SYCL buffer with the type specified by ReinterpretT, dimensions specified by ReinterpretDim and range specified by reinterpretRange.

Returns (2): A reinterpreted SYCL buffer with the type specified by ReinterpretT and dimensions specified by ReinterpretDim.

Throws (1): An exception with the errc::invalid error code if the total size in bytes represented by the type and range of the reinterpreted SYCL buffer (or sub-buffer) does not equal the total size in bytes represented by the type and range of this SYCL buffer (or sub-buffer).

Throws (2): An exception with the errc::invalid error code if the total size in bytes represented by this SYCL buffer (or sub-buffer) is not evenly divisible by sizeof(ReinterpretT).


4.7.2.3. Properties

This section describes the properties that can be passed in the propList parameter of the buffer constructors.


property::buffer::use_host_ptr
namespace sycl::property::buffer {
class use_host_ptr {
  use_host_ptr();  (1)
};
} // namespace sycl::property::buffer

The use_host_ptr property adds the requirement that the SYCL runtime must not allocate any memory for the SYCL buffer and instead uses the provided host pointer directly. This prevents the SYCL runtime from allocating additional temporary storage on the host.

This property has a special guarantee for buffers that are constructed from a hostData pointer. If a host_accessor is constructed from such a buffer, then the address of the reference type returned from the accessor’s member functions such as operator[](id<>) will be the same as the corresponding hostData address.

Effects (1): Constructs an use_host_ptr property object.


property::buffer::use_mutex
namespace sycl::property::buffer {
class use_mutex {
  use_mutex(std::mutex& mutexRef);

  std::mutex* get_mutex_ptr() const;
};
} // namespace sycl::property::buffer

The use_mutex property is valid for the SYCL buffer, unsampled_image and sampled_image classes. The property adds the requirement that the memory which is owned by the SYCL buffer can be shared with the application via a std::mutex provided to the property. The mutex m is locked by the runtime whenever the data is in use and unlocked otherwise. The contents of hostData are guaranteed to reflect the contents of the buffer when the std::mutex is unlocked by the runtime.


property::buffer::use_mutex constructor
use_mutex(std::mutex& mutexRef);

Effects: Constructs a SYCL use_mutex property instance with a reference to mutexRef.


property::buffer::use_mutex::get_mutex_ptr
std::mutex* get_mutex_ptr() const;

Returns: A pointer to the std::mutex provided when constructing this property.


property::buffer::context_bound
namespace sycl::property::buffer {
class context_bound {
 public:
  context_bound(context boundContext);

  context get_context() const;
};
} // namespace sycl::property::buffer

The context_bound property adds the requirement that the SYCL buffer can only be associated with a single SYCL context that is provided to the property.


property::buffer::context_bound constructor
context_bound(context boundContext);

Effects: Constructs a SYCL context_bound property instance with a copy of a SYCL context.


property::buffer::context_bound::get_context
context get_context() const;

Returns: The context provided when constructing this property.


4.7.2.4. Destruction rules

Buffers are reference-counted. When a buffer value is constructed from another buffer, the two values reference the same buffer and a reference count is incremented. When a buffer value is destroyed, the reference count is decremented. Only when there are no more buffer values that reference a specific buffer is the actual buffer destroyed and the buffer destruction behavior defined below is followed.

If any error occurs on buffer destruction, it is reported via the associated queue’s asynchronous error handling mechanism.

The basic rule for the blocking behavior of a buffer destructor is that it blocks if there is some data to write back because a write accessor on it has been created, or if the buffer was constructed with attached host memory and is still in use.

More precisely:

  1. A buffer can be constructed from a range (and without a hostData pointer). The memory management for this type of buffer is entirely handled by the SYCL system. The destructor for this type of buffer does not need to block, even if work on the buffer has not completed. Instead, the SYCL system frees any storage required for the buffer asynchronously when it is no longer in use in queues. The initial contents of the buffer are unspecified.

  2. A buffer can be constructed from a hostData pointer. The buffer will use this host memory for its full lifetime, but the contents of this host memory are unspecified for the lifetime of the buffer. If the host memory is modified on the host or if it is used to construct another buffer or image during the lifetime of this buffer, then the results are undefined. The initial contents of the buffer will be the contents of the host memory at the time of construction.

    When the buffer is destroyed, the destructor will block until all work in queues on the buffer have completed, then copy the contents of the buffer back to the host memory (if required) and then return.

    1. If the type of the host data is const, then the buffer is read-only; only read accessors are allowed on the buffer and no-copy-back to host memory is performed (although the host memory must still be kept available for use by SYCL). When using the default buffer allocator, the const-ness of the type will be removed in order to allow host allocation of memory, which will allow temporary host copies of the data by the SYCL runtime, for example for speeding up host accesses.

      When the buffer is destroyed, the destructor will block until all work in queues on the buffer have completed and then return, as there is no copy of data back to host.

    2. If the type of the host data is not const but the pointer to host data is const, then the read-only restriction applies only on host and not on device accesses.

      When the buffer is destroyed, the destructor will block until all work in queues on the buffer have completed.

  3. A buffer can be constructed using a shared_ptr to host data. This pointer is shared between the SYCL application and the runtime. In order to allow synchronization between the application and the runtime a mutex is used which will be locked by the runtime whenever the data is in use, and unlocked when it is no longer needed.

    The shared_ptr reference counting is used in order to prevent destroying the buffer host data prematurely. If the shared_ptr is deleted from the user application before buffer destruction, the buffer can continue securely because the pointer hasn’t been destroyed yet. It will not copy data back to the host before destruction, however, as the application side has already deleted its copy.

    Note that since there is an implicit conversion of a std::unique_ptr to a std::shared_ptr, a std::unique_ptr can also be used to pass the ownership to the SYCL runtime.

  4. A buffer can be constructed from a pair of iterator values. In this case, the buffer construction will copy the data from the data range defined by the iterator pair. The destructor will not copy back any data and does not need to block.

  5. A buffer can be constructed from a container on which std::data(container) and std::size(container) are well-formed. The initial contents of the buffer will be the contents of the container at the time of construction.

    The buffer may use the memory within the container for its full lifetime, and the contents of this memory are unspecified for the lifetime of the buffer. If the container memory is modified by the host during the lifetime of this buffer, then the results are undefined.

    When the buffer is destroyed, the destructor will block until all work in queues on the buffer have completed. If the return type of std::data(container) is not const then the destructor will also copy the contents of the buffer to the container (if required).

If set_final_data() is used to change where to write the data back to, then the destructor of the buffer will block if a write accessor on it has been created.

A sub-buffer object can be created which is a sub-range reference to a base buffer. This sub-buffer can be used to create accessors to the base buffer, which have access to the range specified at time of construction of the sub-buffer. Sub-buffers cannot be created from sub-buffers, but only from a base buffer which is not already a sub-buffer.

Sub-buffers must be constructed from a contiguous region of memory in a buffer. This requirement is potentially non-intuitive when working with buffers that have dimensionality larger than one, but maps to one-dimensional SYCL backend native allocations without performance cost due to index mapping computation. For example:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
buffer<int, 2> parent_buffer{
    range<2>{8, 8}};  // Create 2-d buffer with 8x8 ints

// OK: Contiguous region from middle of buffer
buffer<int, 2> sub_buf1{parent_buffer, /*offset*/ range<2>{2, 0},
                        /*size*/ range<2>{2, 8}};

// invalid exception: Non-contiguous regions of 2-d buffer
buffer<int, 2> sub_buf2{parent_buffer, /*offset*/ range<2>{2, 0},
                        /*size*/ range<2>{2, 2}};
buffer<int, 2> sub_buf3{parent_buffer, /*offset*/ range<2>{2, 2},
                        /*size*/ range<2>{2, 6}};

// invalid exception: Out-of-bounds size
buffer<int, 2> sub_buf4{parent_buffer, /*offset*/ range<2>{2, 2},
                        /*size*/ range<2>{2, 8}};

4.7.3. Images

The classes unsampled_image (Table 16) and sampled_image (Table 18) define shared image data of one, two or three dimensions, that can be used by kernels in queues and have to be accessed using the image accessor classes.

The constructors and member functions of the SYCL unsampled_image and sampled_image class templates are listed in Table 16, Table 17, Table 18 and Table 19, respectively. The additional common special member functions and common member functions are listed in Table 7 and Table 8, respectively.

Where relevant, it is the responsibility of the user to ensure that the format of the data matches the format described by image_format.

The allocator template parameter of the SYCL unsampled_image and sampled_image classes can be any allocator type including a custom allocator, however it must allocate in units of std::byte.

For any image that is constructed with the range with an element type size in bytes of s, the image row pitch and image slice pitch should be calculated as follows:

The SYCL unsampled_image and sampled_image class templates provide the common reference semantics (see Section 4.5.2).

4.7.3.1. Unsampled image interface

Each constructor of the unsampled_image takes an image_format to describe the data layout of the image data.

Each constructor additionally takes as the last parameter an optional SYCL property_list to provide properties to the SYCL unsampled_image.

The SYCL unsampled_image class template takes a template parameter AllocatorT for specifying an allocator which is used by the SYCL runtime when allocating temporary memory on the host. If no template argument is provided, the default allocator for the SYCL unsampled_image class image_allocator is used (see Section 4.7.1.1).

  1
  2
  3
  4
  5
  6
  7
  8
  9
 10
 11
 12
 13
 14
 15
 16
 17
 18
 19
 20
 21
 22
 23
 24
 25
 26
 27
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
namespace sycl {

enum class image_format : /* unspecified */ {
  r8g8b8a8_unorm,
  r16g16b16a16_unorm,
  r8g8b8a8_sint,
  r16g16b16a16_sint,
  r32b32g32a32_sint,
  r8g8b8a8_uint,
  r16g16b16a16_uint,
  r32b32g32a32_uint,
  r16b16g16a16_sfloat,
  r32g32b32a32_sfloat,
  b8g8r8a8_unorm
};

template <int Dimensions = 1, typename AllocatorT = sycl::image_allocator>
class unsampled_image {
 public:
  unsampled_image(image_format format, const range<Dimensions>& rangeRef,
                  const property_list& propList = {});

  unsampled_image(image_format format, const range<Dimensions>& rangeRef,
                  AllocatorT allocator, const property_list& propList = {});

  /* Available only when: Dimensions > 1 */
  unsampled_image(image_format format, const range<Dimensions>& rangeRef,
                  const range<Dimensions - 1>& pitch,
                  const property_list& propList = {});

  /* Available only when: Dimensions > 1 */
  unsampled_image(image_format format, const range<Dimensions>& rangeRef,
                  const range<Dimensions - 1>& pitch, AllocatorT allocator,
                  const property_list& propList = {});

  unsampled_image(void* hostPointer, image_format format,
                  const range<Dimensions>& rangeRef,
                  const property_list& propList = {});

  unsampled_image(void* hostPointer, image_format format,
                  const range<Dimensions>& rangeRef, AllocatorT allocator,
                  const property_list& propList = {});

  /* Available only when: Dimensions > 1 */
  unsampled_image(void* hostPointer, image_format format,
                  const range<Dimensions>& rangeRef,
                  const range<Dimensions - 1>& pitch,
                  const property_list& propList = {});

  /* Available only when: Dimensions > 1 */
  unsampled_image(void* hostPointer, image_format format,
                  const range<Dimensions>& rangeRef,
                  const range<Dimensions - 1>& pitch, AllocatorT allocator,
                  const property_list& propList = {});

  unsampled_image(std::shared_ptr<void>& hostPointer, image_format format,
                  const range<Dimensions>& rangeRef,
                  const property_list& propList = {});

  unsampled_image(std::shared_ptr<void>& hostPointer, image_format format,
                  const range<Dimensions>& rangeRef, AllocatorT allocator,
                  const property_list& propList = {});

  /* Available only when: Dimensions > 1 */
  unsampled_image(std::shared_ptr<void>& hostPointer, image_format format,
                  const range<Dimensions>& rangeRef,
                  const range<Dimensions - 1>& pitch,
                  const property_list& propList = {});

  /* Available only when: Dimensions > 1 */
  unsampled_image(std::shared_ptr<void>& hostPointer, image_format format,
                  const range<Dimensions>& rangeRef,
                  const range<Dimensions - 1>& pitch, AllocatorT allocator,
                  const property_list& propList = {});

  /* -- common interface members -- */

  /* -- property interface members -- */

  range<Dimensions> get_range() const;

  /* Available only when: Dimensions > 1 */
  range<Dimensions - 1> get_pitch() const;

  std::size_t byte_size() const noexcept;

  std::size_t size() const noexcept;

  AllocatorT get_allocator() const;

  template <typename DataT,
            access_mode Mode = (std::is_const_v<DataT>
                                    ? access_mode::read
                                    : access_mode::read_write),
            image_target Targ = image_target::device>
  unsampled_image_accessor<DataT, Dimensions, Mode, Targ>
  get_access(handler& commandGroupHandler, const property_list& propList = {});

  template <typename DataT, access_mode Mode = (std::is_const_v<DataT>
                                                    ? access_mode::read
                                                    : access_mode::read_write)>
  host_unsampled_image_accessor<DataT, Dimensions, Mode>
  get_host_access(const property_list& propList = {});

  template <typename Destination = std::nullptr_t>
  void set_final_data(Destination finalData = nullptr);

  void set_write_back(bool flag = true);
};

} // namespace sycl
Table 16. Constructors of the unsampled_image class template
Constructor Description
unsampled_image(image_format format,
                const range<Dimensions>& rangeRef,
                const property_list& propList = {})

Construct a SYCL unsampled_image instance with uninitialized memory. The constructed SYCL unsampled_image will use a default constructed AllocatorT when allocating memory on the host. The element size of the constructed SYCL unsampled_image will be derived from the format parameter. The range of the constructed SYCL unsampled_image is specified by the rangeRef parameter provided. The pitch of the constructed SYCL unsampled_image will be the default size determined by the SYCL runtime. Unless the member function set_final_data() is called with a valid non-null pointer, there will be no write back on destruction. Zero or more properties can be provided to the constructed SYCL unsampled_image via an instance of property_list.

unsampled_image(image_format format,
                const range<Dimensions>& rangeRef,
                AllocatorT allocator,
                const property_list& propList = {})

Construct a SYCL unsampled_image instance with uninitialized memory. The constructed SYCL unsampled_image will use the allocator parameter provided when allocating memory on the host. The element size of the constructed SYCL unsampled_image will be derived from the format parameter. The range of the constructed SYCL unsampled_image is specified by the rangeRef parameter provided. The pitch of the constructed SYCL unsampled_image will be the default size determined by the SYCL runtime. Unless the member function set_final_data() is called with a valid non-null pointer, there will be no write back on destruction. Zero or more properties can be provided to the constructed SYCL unsampled_image via an instance of property_list.

unsampled_image(image_format format,
                const range<Dimensions>& rangeRef,
                const range<Dimensions - 1>& pitch,
                const property_list& propList = {})

Available only when: Dimensions > 1.

Construct a SYCL unsampled_image instance with uninitialized memory. The constructed SYCL unsampled_image will use a default constructed AllocatorT when allocating memory on the host. The element size of the constructed SYCL unsampled_image will be derived from the format parameter. The range of the constructed SYCL unsampled_image is specified by the rangeRef parameter provided. The pitch of the constructed SYCL unsampled_image will be the pitch parameter provided. Unless the member function set_final_data() is called with a valid non-null pointer, there will be no write back on destruction. Zero or more properties can be provided to the constructed SYCL unsampled_image via an instance of property_list.

unsampled_image(image_format format,
                const range<Dimensions>& rangeRef,
                const range<Dimensions - 1>& pitch,
                AllocatorT allocator,
                const property_list& propList = {})

Available only when: Dimensions > 1.

Construct a SYCL unsampled_image instance with uninitialized memory. The constructed SYCL unsampled_image will use the allocator parameter provided when allocating memory on the host. The element size of the constructed SYCL unsampled_image will be derived from the format parameter. The range of the constructed SYCL unsampled_image is specified by the rangeRef parameter provided. The pitch of the constructed SYCL unsampled_image will be the pitch parameter provided. Unless the member function set_final_data() is called with a valid non-null pointer, there will be no write back on destruction. Zero or more properties can be provided to the constructed SYCL unsampled_image via an instance of property_list.

unsampled_image(void* hostPointer, image_format format,
                const range<Dimensions>& rangeRef,
                const property_list& propList = {})

Construct a SYCL unsampled_image instance with the hostPointer parameter provided. The unsampled_image assumes exclusive access to this memory for the duration of its lifetime. The constructed SYCL unsampled_image will use a default constructed AllocatorT when allocating memory on the host. The element size of the constructed SYCL unsampled_image will be derived from the format parameter. The range of the constructed SYCL unsampled_image is specified by the rangeRef parameter provided. The pitch of the constructed SYCL unsampled_image will be the default size determined by the SYCL runtime. Unless the member function set_final_data() is called with a valid non-null pointer, any memory allocated by the SYCL runtime is written back to hostPointer. Zero or more properties can be provided to the constructed SYCL unsampled_image via an instance of property_list.

unsampled_image(void* hostPointer, image_format format,
                const range<Dimensions>& rangeRef,
                AllocatorT allocator,
                const property_list& propList = {})

Construct a SYCL unsampled_image instance with the hostPointer parameter provided. The unsampled_image assumes exclusive access to this memory for the duration of its lifetime. The constructed SYCL unsampled_image will use the allocator parameter provided when allocating memory on the host. The element size of the constructed SYCL unsampled_image will be derived from the format parameter. The range of the constructed SYCL unsampled_image is specified by the rangeRef parameter provided. The pitch of the constructed SYCL unsampled_image will be the default size determined by the SYCL runtime. Unless the member function set_final_data() is called with a valid non-null pointer, any memory allocated by the SYCL runtime is written back to hostPointer. Zero or more properties can be provided to the constructed SYCL unsampled_image via an instance of property_list.

unsampled_image(void* hostPointer, image_format format,
                const range<Dimensions>& rangeRef,
                const range<Dimensions - 1>& pitch,
                const property_list& propList = {})

Available only when: Dimensions > 1

Construct a SYCL unsampled_image instance with the hostPointer parameter provided. The unsampled_image assumes exclusive access to this memory for the duration of its lifetime. The constructed SYCL unsampled_image will use a default constructed AllocatorT when allocating memory on the host. The element size of the constructed SYCL unsampled_image will be derived from the format parameter. The range of the constructed SYCL unsampled_image is specified by the rangeRef parameter provided. The pitch of the constructed SYCL unsampled_image will be the pitch parameter provided. Unless the member function set_final_data() is called with a valid non-null pointer, any memory allocated by the SYCL runtime is written back to hostPointer. Zero or more properties can be provided to the constructed SYCL unsampled_image via an instance of property_list.

unsampled_image(void* hostPointer, image_format format,
                const range<Dimensions>& rangeRef,
                const range<Dimensions - 1>& pitch,
                AllocatorT allocator,
                const property_list& propList = {})

Available only when: Dimensions > 1.

Construct a SYCL unsampled_image instance with the hostPointer parameter provided. The unsampled_image assumes exclusive access to this memory for the duration of its lifetime. The constructed SYCL unsampled_image will use the allocator parameter provided when allocating memory on the host. The element size of the constructed SYCL unsampled_image will be derived from the format parameter. The range of the constructed SYCL unsampled_image is specified by the rangeRef parameter provided. The pitch of the constructed SYCL unsampled_image will be the pitch parameter provided. Unless the member function set_final_data() is called with a valid non-null pointer, any memory allocated by the SYCL runtime is written back to hostPointer. Zero or more properties can be provided to the constructed SYCL unsampled_image via an instance of property_list.

unsampled_image(std::shared_ptr<void>& hostPointer,
                image_format format,
                const range<Dimensions>& rangeRef,
                const property_list& propList = {})

When hostPointer is not empty, construct a SYCL unsampled_image with the contents of its stored pointer. The unsampled_image assumes exclusive access to this memory for the duration of its lifetime. The unsampled_image also creates its own internal copy of the shared_ptr that shares ownership of the hostData memory, which means the application can safely release ownership of this shared_ptr when the constructor returns.

When hostPointer is empty, construct a SYCL unsampled_image with uninitialized memory.

The constructed SYCL unsampled_image will use a default constructed AllocatorT when allocating memory on the host. The element size of the constructed SYCL unsampled_image will be derived from the format parameter. The range of the constructed SYCL unsampled_image is specified by the rangeRef parameter provided. The pitch of the constructed SYCL unsampled_image will be the default size determined by the SYCL runtime. Unless the member function set_final_data() is called with a valid non-null pointer, any memory allocated by the SYCL runtime is written back to hostPointer. Zero or more properties can be provided to the constructed SYCL unsampled_image via an instance of property_list.

unsampled_image(std::shared_ptr<void>& hostPointer,
                image_format format,
                const range<Dimensions>& rangeRef,
                AllocatorT allocator,
                const property_list& propList = {})

When hostPointer is not empty, construct a SYCL unsampled_image with the contents of its stored pointer. The unsampled_image assumes exclusive access to this memory for the duration of its lifetime. The unsampled_image also creates its own internal copy of the shared_ptr that shares ownership of the hostData memory, which means the application can safely release ownership of this shared_ptr when the constructor returns.

When hostPointer is empty, construct a SYCL unsampled_image with uninitialized memory.

The constructed SYCL unsampled_image will use the allocator parameter provided when allocating memory on the host. The element size of the constructed SYCL unsampled_image will be derived from the format parameter. The range of the constructed SYCL unsampled_image is specified by the rangeRef parameter provided. The pitch of the constructed SYCL unsampled_image will be the default size determined by the SYCL runtime. Unless the member function set_final_data() is called with a valid non-null pointer, any memory allocated by the SYCL runtime is written back to hostPointer. Zero or more properties can be provided to the constructed SYCL unsampled_image via an instance of property_list.

unsampled_image(std::shared_ptr<void>& hostPointer,
                image_format format,
                const range<Dimensions>& rangeRef,
                const range<Dimensions - 1>& pitch,
                const property_list& propList = {})

When hostPointer is not empty, construct a SYCL unsampled_image with the contents of its stored pointer. The unsampled_image assumes exclusive access to this memory for the duration of its lifetime. The unsampled_image also creates its own internal copy of the shared_ptr that shares ownership of the hostData memory, which means the application can safely release ownership of this shared_ptr when the constructor returns.

When hostPointer is empty, construct a SYCL unsampled_image with uninitialized memory.

The constructed SYCL unsampled_image will use a default constructed AllocatorT when allocating memory on the host. The element size of the constructed SYCL unsampled_image will be derived from the format parameter. The range of the constructed SYCL unsampled_image is specified by the rangeRef parameter provided. The pitch of the constructed SYCL unsampled_image will be the pitch parameter provided. Unless the member function set_final_data() is called with a valid non-null pointer, any memory allocated by the SYCL runtime is written back to hostPointer. Zero or more properties can be provided to the constructed SYCL unsampled_image via an instance of property_list.

unsampled_image(std::shared_ptr<void>& hostPointer,
                image_format format,
                const range<Dimensions>& rangeRef,
                const range<Dimensions - 1>& pitch,
                AllocatorT allocator,
                const property_list& propList = {})

When hostPointer is not empty, construct a SYCL unsampled_image with the contents of its stored pointer. The unsampled_image assumes exclusive access to this memory for the duration of its lifetime. The unsampled_image also creates its own internal copy of the shared_ptr that shares ownership of the hostData memory, which means the application can safely release ownership of this shared_ptr when the constructor returns.

When hostPointer is empty, construct a SYCL unsampled_image with uninitialized memory.

The constructed SYCL unsampled_image will use the allocator parameter provided when allocating memory on the host. The element size of the constructed SYCL unsampled_image will be derived from the format parameter. The range of the constructed SYCL unsampled_image is specified by the rangeRef parameter provided. The pitch of the constructed SYCL unsampled_image will be the pitch parameter provided. Unless the member function set_final_data() is called with a valid non-null pointer, any memory allocated by the SYCL runtime is written back to hostPointer. Zero or more properties can be provided to the constructed SYCL unsampled_image via an instance of property_list.

Table 17. Member functions of the unsampled_image class template
Member function Description
range<Dimensions> get_range() const

Return a range object representing the size of the image in terms of the number of elements in each dimension as passed to the constructor.

range<Dimensions - 1> get_pitch() const

Available only when: Dimensions > 1.

Return a range object representing the pitch of the image in bytes.

std::size_t size() const noexcept

Returns the total number of elements in the image. Equal to get_range()[0] * ... * get_range()[Dimensions-1].

std::size_t byte_size() const noexcept

Returns the size of the image storage in bytes. The number of bytes may be greater than size()*element size due to padding of elements, rows and slices of the image for efficient access.

AllocatorT get_allocator() const

Returns the allocator provided to the image.

template <typename DataT,
         access_mode Mode = (std::is_const_v<DataT>
                                 ? access_mode::read
                                 : access_mode::read_write),
         image_target Targ = image_target::device>
unsampled_image_accessor<DataT, Dimensions, Mode, Targ>
get_access(handler& commandGroupHandler)

Returns a valid unsampled_image_accessor to the unsampled image with the specified data type, access mode and target in the command group.

template <typename DataT, access_mode Mode = (std::is_const_v<DataT>
                                                   ? access_mode::read
                                                   : access_mode::read_write)>
host_unsampled_image_accessor<DataT, Dimensions, Mode> get_host_access();

Returns a valid host_unsampled_image_accessor to the unsampled image with the specified data type and access mode.

template <typename Destination = std::nullptr_t>
void set_final_data(Destination finalData = nullptr)

The finalData point to where the output of all the image processing is going to be copied to at destruction time, if the image was involved with a write accessor.

Destination can be either an output iterator, or a std::weak_ptr<T>.

Note that a raw pointer is a special case of output iterator and thus defines the host memory to which the result is to be copied.

In the case of a weak pointer, the output is not copied if the weak pointer has expired.

If Destination is std::nullptr_t, then the copy back will not happen.

void set_write_back(bool flag = true)

This member function allows dynamically forcing or canceling the write-back of the data of an image on destruction according to the value of flag.

Forcing the write-back is similar to what happens during a normal write-back as described in Section 4.7.3.4 and Section 4.7.4.

If there is nowhere to write-back, using this function does not have any effect.

4.7.3.2. Sampled image interface

Each constructor of the sampled_image class requires a pointer to the host data the image will sample, an image_format to describe the data layout and an image_sampler (Section 4.7.8) to describe how to sample the image data.

Each constructor additionally takes as the last parameter an optional SYCL property_list to provide properties to the SYCL sampled_image.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
namespace sycl {

enum class image_format : /* unspecified */ {
  r8g8b8a8_unorm,
  r16g16b16a16_unorm,
  r8g8b8a8_sint,
  r16g16b16a16_sint,
  r32b32g32a32_sint,
  r8g8b8a8_uint,
  r16g16b16a16_uint,
  r32b32g32a32_uint,
  r16b16g16a16_sfloat,
  r32g32b32a32_sfloat,
  b8g8r8a8_unorm
};

template <int Dimensions = 1, typename AllocatorT = sycl::image_allocator>
class sampled_image {
 public:
  sampled_image(const void* hostPointer, image_format format,
                image_sampler sampler, const range<Dimensions>& rangeRef,
                const property_list& propList = {});

  /* Available only when: Dimensions > 1 */
  sampled_image(const void* hostPointer, image_format format,
                image_sampler sampler, const range<Dimensions>& rangeRef,
                const range<Dimensions - 1>& pitch,
                const property_list& propList = {});

  sampled_image(std::shared_ptr<const void>& hostPointer, image_format format,
                image_sampler sampler, const range<Dimensions>& rangeRef,
                const property_list& propList = {});

  /* Available only when: Dimensions > 1 */
  sampled_image(std::shared_ptr<const void>& hostPointer, image_format format,
                image_sampler sampler, const range<Dimensions>& rangeRef,
                const range<Dimensions - 1>& pitch,
                const property_list& propList = {});

  /* -- common interface members -- */

  /* -- property interface members -- */

  range<Dimensions> get_range() const;

  /* Available only when: Dimensions > 1 */
  range<Dimensions - 1> get_pitch() const;

  std::size_t byte_size() const noexcept;

  std::size_t size() const noexcept;

  template <typename DataT, image_target Targ = image_target::device>
  sampled_image_accessor<DataT, Dimensions, Targ>
  get_access(handler& commandGroupHandler, const property_list& propList = {});

  template <typename DataT>
  host_sampled_image_accessor<DataT, Dimensions>
  get_host_access(const property_list& propList = {});
};

} // namespace sycl
Table 18. Constructors of the sampled_image class template
Constructor Description
sampled_image(const void* hostPointer, image_format format,
              image_sampler sampler,
              const range<Dimensions>& rangeRef,
              const property_list& propList = {})

Construct a SYCL sampled_image instance with the hostPointer parameter provided. The sampled_image assumes exclusive access to this memory for the duration of its lifetime. The host address is const, so the host accesses must be read-only. Since, the hostPointer is const, this image is only initialized with this memory and there is no write after its destruction. The element size of the constructed SYCL sampled_image will be derived from the format parameter. Accessors that read the constructed SYCL sampled_image will use the sampler parameter to sample the image. The range of the constructed SYCL sampled_image is specified by the rangeRef parameter provided. The pitch of the constructed SYCL sampled_image will be the default size determined by the SYCL runtime. Zero or more properties can be provided to the constructed SYCL sampled_image via an instance of property_list.

sampled_image(const void* hostPointer, image_format format,
              image_sampler sampler,
              const range<Dimensions>& rangeRef,
              const range<Dimensions - 1>& pitch,
              const property_list& propList = {})

Available only when: Dimensions > 1.

Construct a SYCL sampled_image instance with the hostPointer parameter provided. The sampled_image assumes exclusive access to this memory for the duration of its lifetime. The host address is const, so the host accesses must be read-only. Since, the hostPointer is const, this image is only initialized with this memory and there is no write after destruction. The element size of the constructed SYCL sampled_image will be derived from the format parameter. Accessors that read the constructed SYCL sampled_image will use the sampler parameter to sample the image. The range of the constructed SYCL sampled_image is specified by the rangeRef parameter provided. The pitch of the constructed SYCL sampled_image will be the pitch parameter provided. Zero or more properties can be provided to the constructed SYCL sampled_image via an instance of property_list.

sampled_image(std::shared_ptr<const void>& hostPointer,
              image_format format,
              image_sampler sampler,
              const range<Dimensions>& rangeRef,
              const property_list& propList = {})

When hostPointer is not empty, construct a SYCL sampled_image with the contents of its stored pointer. The sampled_image assumes exclusive access to this memory for the duration of its lifetime. The sampled_image also creates its own internal copy of the shared_ptr that shares ownership of the hostData memory, which means the application can safely release ownership of this shared_ptr when the constructor returns.

When hostPointer is empty, construct a SYCL sampled_image with uninitialized memory.

The host address is const, so the host accesses must be read-only. Since, the hostPointer is const, this image is only initialized with this memory and there is no write after its destruction. The element size of the constructed SYCL sampled_image will be derived from the format parameter. Accessors that read the constructed SYCL sampled_image will use the sampler parameter to sample the image. The range of the constructed SYCL sampled_image is specified by the rangeRef parameter provided. The pitch of the constructed SYCL sampled_image will be the default size determined by the SYCL runtime. Zero or more properties can be provided to the constructed SYCL sampled_image via an instance of property_list.

sampled_image(std::shared_ptr<const void>& hostPointer,
              image_format format,
              image_sampler sampler,
              const range<Dimensions>& rangeRef,
              const range<Dimensions - 1>& pitch,
              const property_list& propList = {})

When hostPointer is not empty, construct a SYCL sampled_image with the contents of its stored pointer. The sampled_image assumes exclusive access to this memory for the duration of its lifetime. The sampled_image also creates its own internal copy of the shared_ptr that shares ownership of the hostData memory, which means the application can safely release ownership of this shared_ptr when the constructor returns.

When hostPointer is empty, construct a SYCL sampled_image with uninitialized memory.

The host address is const, so the host accesses can be read-only. Since, the hostPointer is const, this image is only initialized with this memory and there is no write after its destruction. The element size of the constructed SYCL sampled_image will be derived from the format parameter. Accessors that read the constructed SYCL sampled_image will use the sampler parameter to sample the image. The range of the constructed SYCL sampled_image is specified by the rangeRef parameter provided. The pitch of the constructed SYCL sampled_image will be the pitch parameter provided. Zero or more properties can be provided to the constructed SYCL sampled_image via an instance of property_list.

Table 19. Member functions of the sampled_image class template
Member function Description
range<Dimensions> get_range() const

Return a range object representing the size of the image in terms of the number of elements in each dimension as passed to the constructor.

range<Dimensions - 1> get_pitch() const

Available only when: Dimensions > 1.

Return a range object representing the pitch of the image in bytes.

std::size_t size() const noexcept

Returns the total number of elements in the image. Equal to get_range()[0] * ... * get_range()[Dimensions-1].

std::size_t byte_size() const noexcept

Returns the size of the image storage in bytes. The number of bytes may be greater than size()*element size due to padding of elements, rows and slices of the image for efficient access.

template <typename DataT, image_target Targ = image_target::device>
sampled_image_accessor<DataT, Dimensions, Targ>
get_access(handler& commandGroupHandler)

Returns a valid sampled_image_accessor to the sampled image with the specified data type and target in the command group.

template <typename DataT>
host_sampled_image_accessor<DataT, Dimensions> get_host_access()

Returns a valid host_sampled_image_accessor to the sampled image with the specified data type in the command group.

4.7.3.3. Image properties

The properties that can be provided when constructing the SYCL unsampled_image and sampled_image classes are described in Table 20.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
namespace sycl {
namespace property {
namespace image {
class use_host_ptr {
 public:
  use_host_ptr() = default;
};

class use_mutex {
 public:
  use_mutex(std::mutex& mutexRef);

  std::mutex* get_mutex_ptr() const;
};

class context_bound {
 public:
  context_bound(context boundContext);

  context get_context() const;
};
} // namespace image
} // namespace property
} // namespace sycl
Table 20. Properties supported by the SYCL image classes
Property Description
property::image::use_host_ptr

The use_host_ptr property adds the requirement that the SYCL runtime must not allocate any memory for the image and instead uses the provided host pointer directly. This prevents the SYCL runtime from allocating additional temporary storage on the host.

property::image::use_mutex

The property adds the requirement that the memory which is owned by the SYCL image can be shared with the application via a std::mutex provided to the property. The std::mutex is locked by the runtime whenever the data is in use and unlocked otherwise. The contents of hostData are guaranteed to reflect the contents of the image when the std::mutex is unlocked by the runtime.

property::image::context_bound

The context_bound property adds the requirement that the SYCL image can only be associated with a single SYCL context that is provided to the property.

The constructors and member functions of the image property classes are listed in Table 21 and Table 22

Table 21. Constructors of the image property classes
Constructor Description
property::image::use_host_ptr::use_host_ptr()

Constructs a SYCL use_host_ptr property instance.

property::image::use_mutex::use_mutex(std::mutex& mutexRef)

Constructs a SYCL use_mutex property instance with a reference to mutexRef parameter provided.

property::image::context_bound::context_bound(context boundContext)

Constructs a SYCL context_bound property instance with a copy of a SYCL context.

Table 22. Member functions of the image property classes
Member function Description
std::mutex* property::image::use_mutex::get_mutex_ptr() const

Returns the std::mutex which was specified when constructing this SYCL use_mutex property.

context property::image::context_bound::get_context() const

Returns the context which was specified when constructing this SYCL context_bound property.

4.7.3.4. Image destruction rules

The rules are similar to those described in Section 4.7.2.4.

For the lifetime of the image object, the associated host memory must be left available to the SYCL runtime and the contents of the associated host memory is unspecified until the image object is destroyed. If an image object value is copied, then only a reference to the underlying image object is copied. The underlying image object is reference-counted. Only after all image value references to the underlying image object have been destroyed is the actual image object itself destroyed.

If an image object is constructed with associated host memory, then its destructor blocks until all operations in all SYCL queues on that image object have completed. Any modifications to the image data will be copied back, if necessary, to the associated host memory. Any errors occurring during destruction are reported to any associated context’s asynchronous error handler. If an image object is constructed with a storage object, then the storage object defines what blocking or copying behavior occurs on image object destruction.

4.7.4. Sharing host memory with the SYCL data management classes

In order to allow the SYCL runtime to do memory management and allow for data dependencies, there are two classes defined, buffer and image. The default behavior for them is that a “raw” pointer is given during the construction of the data management class, with full ownership to use it until the destruction of the SYCL object.

In this section we go in greater detail on sharing or explicitly not sharing host memory with the SYCL data classes, and we will use the buffer class as an example. The same rules will apply to images as well.

4.7.4.1. Default behavior

When using a SYCL buffer, the ownership of the pointer passed to the constructor of the class is, by default, passed to SYCL runtime, and that pointer cannot be used on the host side until the buffer or image is destroyed. A SYCL application can access the contents of the memory managed by a SYCL buffer by using a host_accessor as defined in Section 4.7.6. However, there is no guarantee that the host accessor will copy data back to the original host address used in its constructor.

The pointer passed in is the one used to copy data back to the host, if needed, before buffer destruction. The memory pointed by host pointer will not be de-allocated by the runtime, and the data is copied back from the device if there is a need for it.

4.7.4.2. SYCL ownership of the host memory

In the case where there is host memory to be used for initialization of data but there is no intention of using that host memory after the buffer is destroyed, then the buffer can take full ownership of that host memory.

When a buffer owns the host pointer there is no copy back, by default. In this situation, the SYCL application may pass a unique pointer to the host data, which will be then used by the runtime internally to initialize the data in the device.

For example, the following could be used:

1
2
3
4
5
6
{
  auto ptr = std::make_unique<int>(-1234);
  buffer<int, 1> b { std::move(ptr), range { 1 } };
  // ptr is not valid anymore.
  // There is nowhere to copy data back
}

However, optionally the buffer::set_final_data() can be set to a std::weak_ptr to enable copying data back, to another host memory address that will be valid when the buffer is destroyed.

1
2
3
4
5
6
7
8
{
  auto ptr = std::make_unique<int>(-42);
  buffer<int, 1> b { std::move(ptr), range { 1 } };
  // ptr is not valid anymore.
  // There is nowhere to copy data back.
  // To get copy back, a location can be specified:
  b.set_final_data(std::weak_ptr<int> { .... })
}
4.7.4.3. Shared SYCL ownership of the host memory

When an instance of std::shared_ptr is passed to the buffer constructor, then the buffer object and the developer’s application share the memory region. If the shared pointer is still used on the application’s side then the data will be copied back from the buffer or image and will be available to the application after the buffer or image is destroyed.

If the shared_ptr is not empty, the contents of the referenced memory are used to initialize the buffer. If the shared_ptr is empty, then the buffer is created with uninitialized memory.

When the buffer is destroyed and the data have potentially been updated, if the number of copies of the shared pointer outside the runtime is 0, there is no user-side shared pointer to read the data. Therefore the data is not copied out, and the buffer destructor does not need to wait for the data processes to be finished, as the outcome is not needed on the application’s side.

This behavior can be overridden using the set_final_data() member function of the buffer class, which will by any means force the buffer destructor to wait until the data is copied to wherever the set_final_data() member function has put the data (or not wait nor copy if set final data is nullptr).

1
2
3
4
5
6
7
8
{
  std::shared_ptr<int> ptr { data };
  {
    buffer<int, 2> b { ptr, { 10, 10 } };
    // update the data
    [...]
  } // Data is copied back because there is an user side shared_ptr
}
1
2
3
4
5
6
7
8
9
{
  std::shared_ptr<int> ptr { data };
  {
    buffer<int, 2> b { ptr, { 10, 10 } };
    // update the data
    [...]
    ptr.reset();
  } // Data is not copied back, there is no user side shared_ptr.
}

4.7.5. Synchronization primitives

To prevent race conditions between accesses to the host memory owned by a buffer in the SYCL runtime (e.g., by accessors) and in host code, it is necessary to use manual synchronization through a host_accessor, or by passing a std::mutex to the buffer constructor through a property.

When a buffer was constructed with a std::mutex property, the SYCL runtime is required to lock the mutex whenever the data is in use by the runtime, and unlock the mutex when the data is not in use by the SYCL runtime.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
{
  std::mutex m;
  auto shD = std::make_shared<int>(42)
  sycl::buffer b { shD, { sycl::property::buffer::use_mutex { m } } };
  {
    std::lock_guard lck { m };

    // User accesses the data
    do_something(shD);

    /* m is unlocked when lck goes out of scope, either at normal ending of this
       block or if an exception is thrown */
  }
}

When the runtime releases the mutex, the user is guaranteed that the data has been copied back through the shared pointer --- unless the final data destination has been changed using the member function set_final_data().

4.7.6. Accessors

Accessors provide three different capabilities: they provide access to the data managed by a buffer or image, they provide access to local memory on a device, and they define the requirements to memory objects which determine the scheduling of kernels (see Section 3.8.1).

A memory object requirement is created when an accessor is constructed, unless the accessor is a placeholder in which case the requirement is created when the accessor is bound to a command by calling handler::require().

There are several different C++ classes that implement accessors:

  • The accessor class provides access to data in a buffer from within a command.

  • The host_accessor class provides access to data in a buffer from host code that is outside of a command. These accessors are typically used in application scope.

  • The local_accessor class provides access to device local memory from within a SYCL kernel function.

  • The unsampled_image_accessor and sampled_image_accessor classes provide access to data in an unsampled_image and sampled_image from within a command.

  • The host_unsampled_image_accessor and host_sampled_image_accessor classes provide access to data in an unsampled_image and sampled_image from host code that is outside of a command. These accessors are typically used in application scope.

Accessor objects must always be constructed in host code, either in command group scope or in application scope. Whether the constructor blocks until data is available depends on the type of accessor. Those accessors which provide access to data within a command do not block. Instead, these accessors define a requirement which influences the scheduling of the command. Those accessors which provide access to data from host code do block until the data is available on the host.

For those accessors which provide access to data within a command, the member functions which access data should only be called from within the command. Programs which call these member functions from outside of the command are ill formed. The sections below describe exactly which member functions fall into this category.

4.7.6.1. Data type

All accessors have a DataT template parameter which specifies the type of each element that the accessor accesses. For accessor and host_accessor, this type must either match the type of each element in the underlying buffer, or it must be a const qualified version of that type.

For the image accessors (unsampled_image_accessor, sampled_image_accessor, host_unsampled_image_accessor, and host_sampled_image_accessor), DataT must be one of:

  • int4 (vec<std::int32_t,4>),

  • uint4 (vec<std::uint32_t,4>),

  • float4 (vec<float,4>), or

  • half4 (vec<half,4>)

For local_accessor see Section 4.7.6.11 for the allowable DataT types.

4.7.6.2. Access modes

Most accessors have an AccessMode template parameter which specifies whether the accessor can read or write the underlying data. This information is used by the runtime when defining the requirements for the associated command, and it tells the runtime whether data needs to be transferred to or from a device before data can be accessed through the accessor.

The access_mode enumeration, shown in Table 23, describes the potential modes of an accessor. However, not all accessor classes support all modes, so see the description of each class for more details.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
namespace sycl {

enum class access_mode : /* unspecified */ {
  read,
  write,
  read_write,
  discard_write,      // Deprecated in SYCL 2020
  discard_read_write, // Deprecated in SYCL 2020
  atomic              // Deprecated in SYCL 2020
};

namespace access {
// The legacy type "access::mode" is deprecated.
using mode = sycl::access_mode;
} // namespace access

} // namespace sycl
Table 23. Enumeration of access modes available to accessors
access_mode Description
access_mode::read

Read-only access.

access_mode::write

Write-only access.

access_mode::read_write

Read and write access.

4.7.6.3. Deduction tags

Some accessor constructors take a DeductionTagT parameter, which is used to deduce template arguments for the constructor’s class. Each of the access modes in Table 23 has an associated tag, but there are additional tags which set other template parameters in addition to the access mode. The synopsis below shows the namespace scope variables that the implementation provides as possible values for the DeductionTagT parameter.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
namespace sycl {

inline constexpr __unspecified__ read_only;
inline constexpr __unspecified__ read_write;
inline constexpr __unspecified__ write_only;
inline constexpr __unspecified__ read_only_host_task;
inline constexpr __unspecified__ read_write_host_task;
inline constexpr __unspecified__ write_only_host_task;

} // namespace sycl

The precise meaning of these tags depends on the specific accessor class that is being constructed, so they are described more fully below in the section that pertains to each of the accessor types.

4.7.6.4. Properties

All accessor constructors accept a property_list parameter, which affects the semantics of the accessor. Table 24 shows the set of all possible accessor properties and tells which properties are allowed when constructing each accessor class.

1
2
3
4
5
6
7
namespace sycl {
namespace property {
struct no_init {};
} // namespace property

inline constexpr property::no_init no_init;
} // namespace sycl
Table 24. Properties supported by accessors
Property Allowed with Description

property::no_init

accessor
host_accessor
unsampled_image_accessor
host_unsampled_image_accessor

This property is useful when an application expects to write new values to all of the accessor’s elements without reading their previous values. The implementation can use this information to avoid copying the accessor’s data in some cases. Following is a more formal description.

This property is allowed only for accessors with access_mode::write or access_mode::read_write access modes. Attempting to construct an access_mode::read accessor with this property causes an exception with the errc::invalid error code to be thrown.

The usage of this property is different depending on whether the accessor’s underlying data type DataT is an implicit-lifetime type (as defined in the C++ core language). If it is an implicit-lifetime type, the accessor implicitly creates objects of that type with indeterminate values. The application is not required to write values to each element of the accessor, but unwritten elements of the accessor’s buffer or image receive indeterminate values, even if those buffer or image elements previously had defined values. If this is a ranged accessor, this applies only to the elements within the accessor’s range. The values of unwritten elements outside of this range are preserved.

If DataT is not an implicit-lifetime type, the accessor merely allocates uninitialized memory, and the application is responsible for constructing objects in that memory (e.g. by calling placement-new). The application must create an object in each element of the accessor unless the corresponding element of the underlying buffer did not previously contain an object. If this is a ranged accessor, this applies only to the elements within the accessor’s range. The content of objects in the buffer outside of this range is preserved.

As stated above, the property::no_init property requires the application to construct an object for each accessor element when the element’s type is not an implicit-lifetime type (except in the case when the corresponding buffer element did not previously contain an object). The reason for this requirement is to avoid the possibility of overwriting a valid object with indeterminate bytes, for example, when a command using the accessor completes. This means that the implementation can unconditionally copy memory from the device back to the host when the command completes, regardless of whether the DataT type is an implicit-lifetime type.

The constructors of the accessor property classes are listed in Table 25.

Table 25. Constructors of the accessor property classes
Constructor Description
property::no_init::no_init()

Constructs a no_init property instance.

4.7.6.5. Read only accessors

Accessors which have an AccessMode template parameter can be declared as read-only by specifying access_mode::read for the template parameter. A read-only accessor provides read-only access to the underlying data and provides a "read" requirement for the memory object when it is constructed.

The DataT template parameter for a read-only accessor can optionally be const qualified, and the semantics of the accessor are unchanged. For example, an accessor declared with const DataT and access_mode::read has the same semantics as an accessor declared with DataT and access_mode::read.

As detailed in the sections below, some accessor types have a default value for AccessMode, which depends on whether the DataT parameter is const qualified. This provides a convenient way to declare a read-only accessor without explicitly specifying the access mode.

A const qualified DataT is only allowed for a read-only accessor. Programs which specify a const qualified DataT and any access mode other than access_mode::read are ill formed, and the implementation must issue a diagnostic in this case.

Each accessor class also provides implicit conversions between the two forms of read-only accessors. This makes it possible, for example, to assign an accessor whose type has const DataT and access_mode::read to an accessor whose type has DataT and access_mode::read, so long as the other template parameters are the same. There is also an implicit conversion from a read-write accessor to either of the forms of a read-only accessor. These implicit conversions are described in detail for each accessor class in the sections that follow.

4.7.6.6. Accessing elements of an accessor

Accessors of type accessor, host_accessor, and local_accessor can have zero, one, two, or three Dimensions. A zero dimension accessor provides access to a single scalar element via an implicit conversion operator to the underlying type of that element and via an overloaded copy/move assignment operators from the underlying type of the element.

One, two, or three dimensional specializations of these accessors provide access to the elements they contain in two ways. The first way is through a subscript operator that takes an instance of an id class which has the same dimensionality as the accessor. The second way is by passing a single std::size_t value to multiple consecutive subscript operators as specified in Section 3.11.2.

In all these cases, the reference to the contained element is of type const DataT& for read-only accessors and of type DataT& for other accessors.

Accessors of all types have a range that defines the set of indices that may be used to access elements. For buffer accessors, this is the range of the underlying buffer, unless it is a ranged accessor in which case the range comes from the accessor’s constructor. For image accessors, this is the range of the underlying image. Local accessors specify the range when the accessor is constructed. Any attempt to access an element via an index that is outside of this range produces undefined behavior.

4.7.6.7. Container interface

Accessors of type accessor, host_accessor, and local_accessor meet the C++ requirement of ReversibleContainer. The exception to this is that only local_accessor owns the underlying data, meaning that its destructor destroys elements and frees the memory. The accessor and host_accessor types don’t destroy any elements or free the memory on destruction. The iterator for the container interface meets the C++ requirement of LegacyRandomAccessIterator and the underlying pointers/references correspond to the address space specified by the accessor type. For multidimensional accessors the iterator linearizes the data according to Section 3.11.1.

4.7.6.8. Ranged accessors

Accessors of type accessor and host_accessor can be constructed from a sub-range of a buffer by providing a range and offset to the constructor. This limits the elements that can be accessed to the specified sub-range, which allows the implementation to perform certain optimizations such as reducing the amount of memory that needs to be copied to or from a device.

If the ranged accessor is multi-dimensional, the sub-range is allowed to describe a region of memory in the underlying buffer that is not contiguous in the linear address space. It is also legal to construct several ranged accessors for the same underlying buffer, either overlapping or non-overlapping.

A ranged accessor still creates a requisite for the entire underlying buffer, even for the portions not within the range. For example, if one command writes through a ranged accessor to one region of a buffer and a second command reads through a ranged accessor from a non-overlapping region of the same buffer, the second command must still be scheduled after the first because the requisites for the two commands are on the entire buffer, not on the sub-ranges of the ranged accessors.

Most of the accessor member functions which provide a reference to the underlying buffer elements are affected by a ranged accessor’s offset and range. For example, calling operator[](0) on a one-dimensional ranged accessor returns a reference to the element at the position specified by the accessor’s offset, which is not necessarily the first element in the buffer. In addition, the accessor’s iterator functions iterate only over the elements that are within the sub-range.

The only exceptions are the get_pointer and get_multi_ptr member functions, which return a pointer to the beginning of the underlying buffer regardless of the accessor’s offset. Applications using these functions must take care to manually add the offset before dereferencing the pointer because accessing an element that is outside of the accessor’s range results in undefined behavior.

There is no change in behavior for ranged accessors with a range of zero. It still creates a requisite for the entire underlying buffer, and an attempt to access an element produces undefined behavior.

4.7.6.9. Buffer accessor for commands

The accessor class provides access to data in a buffer in three different ways. It can be used to access the buffer’s data from within a SYCL kernel function via the device’s global memory. It can also be used to access the buffer’s data on host from within a host task. Finally, it can be used to get a native backend handle to the buffer from within a host task. The AccessTarget template parameter helps distinguish these three cases as shown in Table 26.

Table 26. Description of access targets for buffer accessors
Access target Meaning

target::device

Access a buffer from a SYCL kernel function via device global memory. Also used to get a native backend handle to the buffer from within a host task.

target::host_task

Access a buffer’s data on host from within a host task.

When an accessor is used from within a SYCL kernel function, the access target must be target::device, target::constant_buffer, or target::local; otherwise the behavior is undefined. See Section 4.7.6.9.4.5 and Section 4.7.6.9.4.7 for a description of the deprecated target::constant_buffer and target::local targets.

When an accessor is used from within a host task, the use of the accessor must correspond to the access target, otherwise the behavior is undefined. If the access target is target::host_task, the accessor may only be used to access the buffer’s data on host, from within the host task function. If the access target is target::device, the accessor may only be used to get a native backend handle for the buffer as described in Section 4.10.2.

The dimensionality of the accessor must match the underlying buffer, however, there is a special case if the buffer is one-dimensional. In this case, the accessor may either be one-dimensional or it may be zero-dimensional. A zero-dimensional accessor has access to just the first element of the buffer, whereas a one-dimensional accessor has access to the entire buffer.

Certain accessor constructors create a "placeholder" accessor. Such an accessor is bound to a buffer and its semantics such as access target and access mode are defined. However, a placeholder accessor is not yet bound to a command group. Before such an accessor can be used in a command, it must be bound by calling handler::require(). Passing a placeholder accessor as an argument to a command without first being bound to a command group with handler::require() will result in undefined behavior.

Implementations are encouraged to throw either a synchronous or an asynchronous exception when a placeholder accessor, that has not been bound to the corresponding command group with handler::require(), is either passed as an argument to or is used inside a command.

4.7.6.9.1. Interface for buffer command accessors

A synopsis of the accessor class is provided below, showing the interface when it is specialized with target::device or target::host_task. Since some of the class types and member functions have the same name and meaning as other accessors, the common types and functions are described in Section 4.7.6.12. The member types are listed in Table 51 and Table 27. The constructors are listed in Table 28, and the member functions are listed in Table 52 and Table 29.

The additional common special member functions and common member functions are listed in Section 4.5.2 in Table 7 and Table 8, respectively. For valid implicit conversions between accessor types refer to Section 4.7.6.9.3. Additionally, accessors of the same type must be equality comparable both in the host application and also in SYCL kernel functions.

  1
  2
  3
  4
  5
  6
  7
  8
  9
 10
 11
 12
 13
 14
 15
 16
 17
 18
 19
 20
 21
 22
 23
 24
 25
 26
 27
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
namespace sycl {

enum class target : /* unspecified */ {
  device,
  host_task,
  constant_buffer,       // Deprecated
  local,                 // Deprecated
  host_buffer,           // Deprecated
  global_buffer = device // Deprecated
};

namespace access {
// The legacy type "access::target" is deprecated.
using sycl::target;

enum class placeholder : /* unspecified */ { // Deprecated
  false_t,
  true_t
};

} // namespace access

template <typename DataT, int Dimensions = 1,
          access_mode AccessMode =
              (std::is_const_v<DataT> ? access_mode::read
                                      : access_mode::read_write),
          target AccessTarget = target::device,
          access::placeholder isPlaceholder = access::placeholder::false_t>
class accessor {
 public:
  using value_type = // const DataT for read-only accessors, DataT otherwise
      __value_type__;
  using reference = value_type&;
  using const_reference = const DataT&;
  template <access::decorated IsDecorated>
  using accessor_ptr =   // multi_ptr to value_type with target address space,
      __pointer_class__; //   unspecified for access_mode::host_task
  using iterator = __unspecified_iterator__<value_type>;
  using const_iterator = __unspecified_iterator__<const value_type>;
  using reverse_iterator = std::reverse_iterator<iterator>;
  using const_reverse_iterator = std::reverse_iterator<const_iterator>;
  using difference_type =
      typename std::iterator_traits<iterator>::difference_type;
  using size_type = std::size_t;

  accessor();

  /* Available only when: (Dimensions == 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, 1, AllocatorT>& bufferRef,
           const property_list& propList = {});

  /* Available only when: (Dimensions == 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, 1, AllocatorT>& bufferRef,
           handler& commandGroupHandlerRef, const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT, typename DeductionTagT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef, DeductionTagT tag,
           const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           handler& commandGroupHandlerRef, const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT, typename DeductionTagT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           handler& commandGroupHandlerRef, DeductionTagT tag,
           const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           range<Dimensions> accessRange, const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT, typename DeductionTagT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           range<Dimensions> accessRange, DeductionTagT tag,
           const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           range<Dimensions> accessRange, id<Dimensions> accessOffset,
           const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT, typename DeductionTagT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           range<Dimensions> accessRange, id<Dimensions> accessOffset, DeductionTagT tag,
           const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           handler& commandGroupHandlerRef, range<Dimensions> accessRange,
           const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT, typename DeductionTagT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           handler& commandGroupHandlerRef, range<Dimensions> accessRange,
           DeductionTagT tag, const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           handler& commandGroupHandlerRef, range<Dimensions> accessRange,
           id<Dimensions> accessOffset, const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT, typename DeductionTagT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           handler& commandGroupHandlerRef, range<Dimensions> accessRange,
           id<Dimensions> accessOffset, DeductionTagT tag,
           const property_list& propList = {});

  /* -- common interface members -- */

  void swap(accessor& other);

  bool is_placeholder() const;

  size_type byte_size() const noexcept;

  size_type size() const noexcept;

  // Deprecated
  size_type max_size() const noexcept;

  // Deprecated
  std::size_t get_size() const;

  // Deprecated
  std::size_t get_count() const;

  bool empty() const noexcept;

  /* Available only when: (Dimensions > 0) */
  range<Dimensions> get_range() const;

  /* Available only when: (Dimensions > 0) */
  id<Dimensions> get_offset() const;

  /* Available only when: (AccessMode != access_mode::atomic && Dimensions == 0) */
  operator reference() const;

  /* Available only when: (AccessMode != access_mode::atomic &&
                           AccessMode != access_mode::read && Dimensions == 0) */
  const accessor& operator=(const value_type& other) const;

  /* Available only when: (AccessMode != access_mode::atomic &&
                           AccessMode != access_mode::read && Dimensions == 0) */
  const accessor& operator=(value_type&& other) const;

  /* Available only when: (Dimensions > 0) */
  reference operator[](id<Dimensions> index) const;

  /* Available only when: (Dimensions > 1) */
  __unspecified__ operator[](std::size_t index) const;

  /* Available only when: (AccessMode != access_mode::atomic && Dimensions == 1)
   */
  reference operator[](std::size_t index) const;

  /* Deprecated
  Available only when: (AccessMode == access_mode::atomic && Dimensions ==  0)
*/
  operator sycl::atomic<DataT, access::address_space::global_space>() const;

  /* Deprecated
  Available only when: (AccessMode == access_mode::atomic && Dimensions == 1) */
  sycl::atomic<DataT, access::address_space::global_space>
  operator[](id<Dimensions> index) const;

  /* Deprecated in SYCL 2020
  Available only when: (AccessTarget == target::device) */
  global_ptr<value_type> get_pointer() const noexcept;

  /* Available only when: (AccessTarget == target::host_task) */
  std::add_pointer_t<value_type> get_pointer() const noexcept;

  /* Available only when: (AccessTarget == target::device) */
  template <access::decorated IsDecorated>
  accessor_ptr<IsDecorated> get_multi_ptr() const noexcept;

  iterator begin() const noexcept;

  iterator end() const noexcept;

  const_iterator cbegin() const noexcept;

  const_iterator cend() const noexcept;

  reverse_iterator rbegin() const noexcept;

  reverse_iterator rend() const noexcept;

  const_reverse_iterator crbegin() const noexcept;

  const_reverse_iterator crend() const noexcept;
};

} // namespace sycl
Table 27. Member types of the accessor class
Member types Description
template <access::decorated IsDecorated> accessor_ptr

If (AccessTarget == target::device): multi_ptr<value_type, access::address_space::global_space, IsDecorated>.

The definition of this type is not specified when (AccessTarget == target::host_task).

Table 28. Constructors of the accessor class
Constructor Description
accessor()

Constructs an empty accessor which fulfills the following post-conditions:

  • (empty() == true)

  • All size queries return 0.

  • The return values of get_pointer() and get_multi_ptr() are unspecified.

  • A default constructed accessor can be passed to a SYCL kernel function, but attempting to access data elements from it produces undefined behavior.

template <typename AllocatorT>
accessor(buffer<DataT, 1, AllocatorT>& bufferRef,
         const property_list& propList = {})

Available only when (Dimensions == 0).

Constructs a placeholder accessor for accessing the first element of a buffer. The optional property_list provides properties for the constructed accessor.

template <typename AllocatorT>
accessor(buffer<DataT, 1, AllocatorT>& bufferRef,
         handler& commandGroupHandlerRef, const property_list& propList = {})

Available only when (Dimensions == 0).

Constructs an accessor for accessing the first element of a buffer within a SYCL kernel function on the queue associated with commandGroupHandlerRef. The optional property_list provides properties for the constructed accessor.

template <typename AllocatorT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs a placeholder accessor for accessing a buffer. The optional property_list provides properties for the constructed accessor.

template <typename AllocatorT, typename DeductionTagT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef, DeductionTagT tag,
         const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs a placeholder accessor for accessing a buffer. The tag is used to deduce template arguments of the accessor as described in Section 4.7.6.9.2. The optional property_list provides properties for the constructed accessor.

template <typename AllocatorT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         handler& commandGroupHandlerRef, const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs an accessor for accessing a buffer within a SYCL kernel function on the queue associated with commandGroupHandlerRef. The optional property_list provides properties for the constructed accessor.

template <typename AllocatorT, typename DeductionTagT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         handler& commandGroupHandlerRef, DeductionTagT tag,
         const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs an accessor for accessing a buffer within a SYCL kernel function on the queue associated with commandGroupHandlerRef. The tag is used to deduce template arguments of the accessor as described in Section 4.7.6.9.2. The optional property_list provides properties for the constructed accessor.

template <typename AllocatorT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         range<Dimensions> accessRange, const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs a placeholder accessor that is a ranged accessor, where the range starts at the beginning of the buffer. The optional property_list provides properties for the constructed accessor.

Throws an exception with the errc::invalid error code if accessRange exceeds the range of bufferRef in any dimension.

template <typename AllocatorT, typename DeductionTagT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         range<Dimensions> accessRange, DeductionTagT tag,
         const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs a placeholder accessor that is a ranged accessor, where the range starts at the beginning of the buffer. The tag is used to deduce template arguments of the accessor as described in Section 4.7.6.9.2. The optional property_list provides properties for the constructed accessor.

Throws an exception with the errc::invalid error code if accessRange exceeds the range of bufferRef in any dimension.

template <typename AllocatorT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         range<Dimensions> accessRange, id<Dimensions> accessOffset,
         const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs a placeholder accessor that is a ranged accessor, where the range starts at an offset from the beginning of the buffer. The optional property_list provides properties for the constructed accessor.

Throws an exception with the errc::invalid error code if the sum of accessRange and accessOffset exceeds the range of bufferRef in any dimension.

template <typename AllocatorT, typename DeductionTagT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         range<Dimensions> accessRange, id<Dimensions> accessOffset, DeductionTagT tag,
         const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs a placeholder accessor that is a ranged accessor, where the range starts at an offset from the beginning of the buffer. The tag is used to deduce template arguments of the accessor as described in Section 4.7.6.9.2. The optional property_list provides properties for the constructed accessor.

Throws an exception with the errc::invalid error code if the sum of accessRange and accessOffset exceeds the range of bufferRef in any dimension.

template <typename AllocatorT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         handler& commandGroupHandlerRef, range<Dimensions> accessRange,
         const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs an accessor that is a ranged accessor, where the range starts at the beginning of the buffer. The accessor can only be used in a SYCL kernel function on the queue associated with commandGroupHandlerRef. The optional property_list provides properties for the constructed accessor.

Throws an exception with the errc::invalid error code if accessRange exceeds the range of bufferRef in any dimension.

template <typename AllocatorT, typename DeductionTagT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         handler& commandGroupHandlerRef, range<Dimensions> accessRange,
         DeductionTagT tag, const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs an accessor that is a ranged accessor, where the range starts at the beginning of the buffer. The accessor can only be used in a SYCL kernel function on the queue associated with commandGroupHandlerRef. The tag is used to deduce template arguments of the accessor as described in Section 4.7.6.9.2. The optional property_list provides properties for the constructed accessor.

Throws an exception with the errc::invalid error code if accessRange exceeds the range of bufferRef in any dimension.

template <typename AllocatorT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         handler& commandGroupHandlerRef, range<Dimensions> accessRange,
         id<Dimensions> accessOffset, const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs an accessor that is a ranged accessor, where the range starts at an offset from the beginning of the buffer. The accessor can only be used in a SYCL kernel function on the queue associated with commandGroupHandlerRef. The optional property_list provides properties for the constructed accessor.

Throws an exception with the errc::invalid error code if the sum of accessRange and accessOffset exceeds the range of bufferRef in any dimension.

template <typename AllocatorT, typename DeductionTagT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         handler& commandGroupHandlerRef, range<Dimensions> accessRange,
         id<Dimensions> accessOffset, DeductionTagT tag,
         const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs an accessor that is a ranged accessor, where the range starts at an offset from the beginning of the buffer. The accessor can only be used in a SYCL kernel function on the queue associated with commandGroupHandlerRef. The tag is used to deduce template arguments of the accessor as described in Section 4.7.6.9.2. The optional property_list provides properties for the constructed accessor.

Throws an exception with the errc::invalid error code if the sum of accessRange and accessOffset exceeds the range of bufferRef in any dimension.

Table 29. Member functions of the accessor class
Member function Description
void swap(accessor& other);

Swaps the contents of the current accessor with the contents of other.

bool is_placeholder() const

Returns true if the accessor is a placeholder. Otherwise returns false.

id<Dimensions> get_offset() const

Available only when (Dimensions > 0).

If this is a ranged accessor, returns the offset that was specified when the accessor was constructed. For other accessors, returns the default constructed id<Dimensions>{}.

global_ptr<value_type> get_pointer() const noexcept

Available only when (AccessTarget == target::device).

Returns a multi_ptr to the start of this accessor’s underlying buffer, even if this is a ranged accessor whose range does not start at the beginning of the buffer. The return value is unspecified if the accessor is empty.

Preconditions: Must be called within a command.

Deprecated in SYCL 2020. Use get_multi_ptr instead.

std::add_pointer_t<value_type> get_pointer() const noexcept

Available only when (AccessTarget == target::host_task).

Returns a pointer to the start of this accessor’s underlying buffer, even if this is a ranged accessor whose range does not start at the beginning of the buffer. The return value is unspecified if the accessor is empty.

Preconditions: Must be called within a command.

template <access::decorated IsDecorated>
accessor_ptr<IsDecorated> get_multi_ptr() const noexcept

Available only when (AccessTarget == target::device).

Returns a multi_ptr to the start of this accessor’s underlying buffer, even if this is a ranged accessor whose range does not start at the beginning of the buffer. The return value is unspecified if the accessor is empty.

Preconditions: Must be called within a command.

const accessor& operator=(const value_type& other) const

Available only when (AccessMode != access_mode::atomic && AccessMode != access_mode::read && Dimensions == 0).

Assignment to the single element that is accessed by this accessor.

Preconditions: Must be called within a command.

const accessor& operator=(value_type&& other) const

Available only when (AccessMode != access_mode::atomic && AccessMode != access_mode::read && Dimensions == 0).

Assignment to the single element that is accessed by this accessor.

Preconditions: Must be called within a command.

4.7.6.9.2. Deduction tags for buffer command accessors

Some accessor constructors take a DeductionTagT parameter, which is used to deduce template arguments. The permissible values for this parameter are listed in Table 30 along with the access mode and accessor target that they imply.

Table 30. Enumeration of tags available for accessor construction
Tag value Access mode Accessor target

read_write

access_mode::read_write

target::device

read_only

access_mode::read

target::device

write_only

access_mode::write

target::device

read_write_host_task

access_mode::read_write

target::host_task

read_only_host_task

access_mode::read

target::host_task

write_only_host_task

access_mode::write

target::host_task

4.7.6.9.3. Read only buffer command accessors and implicit conversions

Table 31 shows the specializations of accessor with target::device or target::host_task that are read-only accessors. There is an implicit conversion between any of these specializations, provided that all other template parameters are the same.

Table 31. Specializations of accessor that are read-only
Data type Access mode

not const-qualified

access_mode::read

const-qualified

access_mode::read

There is also an implicit conversion from the read-write specialization shown in Table 32 to any of the read-only specializations shown in Table 31, provided that all other template parameters are the same.

Table 32. Specializations of accessor that are read-write
Data type Access mode

not const-qualified

access_mode::read_write

4.7.6.9.4. Deprecated features of the accessor class

All of the features defined in this section are deprecated and will likely be removed from a future version of the specification.

4.7.6.9.4.1. Aliased names

The enumerated value target::global_buffer is an alias for target:::device. It has the same type and value as its alias.

The enumerated type access::target is an alias for target, and the enumerated type access::mode is an alias for access_mode.

4.7.6.9.4.2. Discard access modes

An accessor instance specialized with access mode access_mode::discard_write has the same behavior as an accessor instance of mode access_mode::write that is constructed with the property property::no_init.

An accessor instance specialized with access mode access_mode::discard_read_write has the same behavior as an accessor instance of mode access_mode::read_write that is constructed with the property property::no_init.

4.7.6.9.4.3. Placeholder template parameter

The accessor template parameter IsPlaceholder is allowed to be specified, but it has no bearing on whether the accessor instance is a placeholder. This is determined solely by the constructor used to create the instance.

The associated type access::placeholder is also deprecated.

4.7.6.9.4.4. Additional member functions for target::device specialization

Specializations of the accessor class with target::device have the additional member functions described in Table 33.

Table 33. Deprecated member functions of the accessor class
Member function Description
std::size_t get_size() const

Returns the same value as byte_size().

std::size_t get_count() const

Returns the same value as size().

4.7.6.9.4.5. Accessor specialization with target::constant_buffer

The accessor class may be specialized with target target::constant_buffer, which results in an accessor that can be used within a SYCL kernel function to access the contents of a buffer through the device’s constant memory.

As with other accessor specializations, the dimensionality must match the underlying buffer, however there is a special case if the buffer is one-dimensional. In this case, the accessor may either be one-dimensional or it may be zero-dimensional. A zero-dimensional accessor has access to just the first element of the buffer, whereas a one-dimensional accessor has access to the entire buffer.

This specialization of accessor is available only for the access mode access_mode::read.

This accessor type can be constructed as a "placeholder" accessor. As with other accessor specializations that are placeholders, handler::require() must be called before passing a placeholder accessor to a command. Passing a placeholder accessor as an argument to a command without first being bound to a command group with handler::require() will result in undefined behavior.

A synopsis for this specialization of accessor is provided below. Since some of the class types and member functions have the same name and meaning as other accessors, the common types and functions are described in Section 4.7.6.9.4.8. The member types are listed in Table 40. The constructors are listed in Table 34, and the member functions are listed in Table 41 and Table 35.

The additional common special member functions and common member functions are listed in Section 4.5.2 in Table 7 and Table 8, respectively. Additionally, accessors of the same type must be equality comparable.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
namespace sycl {

template <typename DataT, int Dimensions, access_mode AccessMode,
          target AccessTarget, access::placeholder IsPlaceholder>
class accessor {
 public:
  using value_type = const DataT;
  using reference = const DataT&;
  using const_reference = const DataT&;

  /* Available only when: (Dimensions == 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, 1, AllocatorT>& bufferRef,
           const property_list& propList = {});

  /* Available only when: (Dimensions == 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, 1, AllocatorT>& bufferRef,
           handler& commandGroupHandlerRef, const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           handler& commandGroupHandlerRef, const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           range<Dimensions> accessRange, const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           range<Dimensions> accessRange, id<Dimensions> accessOffset,
           const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           handler& commandGroupHandlerRef, range<Dimensions> accessRange,
           const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           handler& commandGroupHandlerRef, range<Dimensions> accessRange,
           id<Dimensions> accessOffset, const property_list& propList = {});

  /* -- common interface members -- */

  bool is_placeholder() const;

  std::size_t get_size() const noexcept;

  std::size_t get_count() const noexcept;

  /* Available only when: (Dimensions > 0) */
  range<Dimensions> get_range() const;

  /* Available only when: (Dimensions > 0) */
  id<Dimensions> get_offset() const;

  /* Available only when: (Dimensions == 0) */
  operator reference() const;

  /* Available only when: (Dimensions > 0) */
  reference operator[](id<Dimensions> index) const;

  /* Available only when: (Dimensions > 1) */
  __unspecified__ operator[](std::size_t index) const;

  /* Available only when: (Dimensions == 1) */
  reference operator[](std::size_t index) const;

  constant_ptr<DataT> get_pointer() const noexcept;
};

} // namespace sycl
Table 34. Constructors of the deprecated constant accessor
Constructor Description
template <typename AllocatorT>
accessor(buffer<DataT, 1, AllocatorT>& bufferRef,
         const property_list& propList = {})

Available only when (Dimensions == 0).

Constructs a placeholder accessor for accessing the first element of a buffer. The optional property_list provides properties for the constructed accessor.

template <typename AllocatorT>
accessor(buffer<DataT, 1, AllocatorT>& bufferRef,
         handler& commandGroupHandlerRef, const property_list& propList = {})

Available only when (Dimensions == 0).

Constructs an accessor for accessing the first element of a buffer within a SYCL kernel function on the queue associated with commandGroupHandlerRef. The optional property_list provides properties for the constructed accessor.

template <typename AllocatorT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs a placeholder accessor for accessing a buffer. The optional property_list provides properties for the constructed accessor.

template <typename AllocatorT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         handler& commandGroupHandlerRef, const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs an accessor for accessing a buffer within a SYCL kernel function on the queue associated with commandGroupHandlerRef. The optional property_list provides properties for the constructed accessor.

template <typename AllocatorT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         range<Dimensions> accessRange, const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs a placeholder accessor that is a ranged accessor, where the range starts at the beginning of the buffer. The optional property_list provides properties for the constructed accessor.

Throws an exception with the errc::invalid error code if accessRange exceeds the range of bufferRef in any dimension.

template <typename AllocatorT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         range<Dimensions> accessRange, id<Dimensions> accessOffset,
         const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs a placeholder accessor that is a ranged accessor, where the range starts at an offset from the beginning of the buffer. The optional property_list provides properties for the constructed accessor.

Throws an exception with the errc::invalid error code if the sum of accessRange and accessOffset exceeds the range of bufferRef in any dimension.

template <typename AllocatorT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         handler& commandGroupHandlerRef, range<Dimensions> accessRange,
         const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs an accessor that is a ranged accessor, where the range starts at the beginning of the buffer. The accessor can only be used in a SYCL kernel function on the queue associated with commandGroupHandlerRef. The optional property_list provides properties for the constructed accessor.

Throws an exception with the errc::invalid error code if accessRange exceeds the range of bufferRef in any dimension.

template <typename AllocatorT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         handler& commandGroupHandlerRef, range<Dimensions> accessRange,
         id<Dimensions> accessOffset, const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs an accessor that is a ranged accessor, where the range starts at an offset from the beginning of the buffer. The accessor can only be used in a SYCL kernel function on the queue associated with commandGroupHandlerRef. The optional property_list provides properties for the constructed accessor.

Throws an exception with the errc::invalid error code if the sum of accessRange and accessOffset exceeds the range of bufferRef in any dimension.

Table 35. Member functions of the deprecated constant accessor
Member function Description
bool is_placeholder() const

Returns true if the accessor was constructed as a placeholder and returns false otherwise.

id<Dimensions> get_offset() const

Available only when (Dimensions > 0).

If this is a ranged accessor, returns the offset that was specified when the accessor was constructed, otherwise returns the default constructed id<Dimensions>{}.

constant_ptr<DataT> get_pointer() const noexcept

Returns a multi_ptr to the start of this accessor’s underlying buffer, even if this is a ranged accessor whose range does not start at the beginning of the buffer. The return value is unspecified if the accessor is empty.

Preconditions: Must be called within a command.

4.7.6.9.4.6. Accessor specialization with target::host_buffer

The accessor class may be specialized with target target::host_buffer, which results in a host accessor similar to host_accessor. This specialization provides access to data in a buffer from host code that is outside of a command, and constructors of this specialization block until the requested data is available on the host.

As with other accessor specializations, the dimensionality must match the underlying buffer, however there is a special case if the buffer is one-dimensional. In this case, the accessor may either be one-dimensional or it may be zero-dimensional. A zero-dimensional accessor has access to just the first element of the buffer, whereas a one-dimensional accessor has access to the entire buffer.

This specialization of accessor is available for all access modes except for access_mode::atomic.

A synopsis for this specialization of accessor is provided below. Since some of the class types and member functions have the same name and meaning as other accessors, the common types and functions are described in Section 4.7.6.9.4.8. The member types are listed in Table 40. The constructors are listed in Table 36, and the member functions are listed in Table 41 and Table 37.

The additional common special member functions and common member functions are listed in Section 4.5.2 in Table 7 and Table 8, respectively. Additionally, accessors of the same type must be equality comparable.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
namespace sycl {

template <typename DataT, int Dimensions, access_mode AccessMode,
          target AccessTarget, access::placeholder IsPlaceholder>
class accessor {
 public:
  using value_type = // const DataT for access_mode::read, DataT otherwise
      __value_type__;
  using reference = value_type&;
  using const_reference = const DataT&;

  /* Available only when: (Dimensions == 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, 1, AllocatorT>& bufferRef,
           const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           range<Dimensions> accessRange, const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT>
  accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
           range<Dimensions> accessRange, id<Dimensions> accessOffset,
           const property_list& propList = {});

  /* -- common interface members -- */

  bool is_placeholder() const;

  std::size_t get_size() const;

  std::size_t get_count() const;

  /* Available only when: (Dimensions > 0) */
  range<Dimensions> get_range() const;

  /* Available only when: (Dimensions > 0) */
  id<Dimensions> get_offset() const;

  /* Available only when: (Dimensions == 0) */
  operator reference() const;

  /* Available only when: (Dimensions > 0) */
  reference operator[](id<Dimensions> index) const;

  /* Available only when: (Dimensions > 1) */
  __unspecified__ operator[](std::size_t index) const;

  /* Available only when: (Dimensions == 1) */
  reference operator[](std::size_t index) const;

  std::add_pointer_t<value_type> get_pointer() const noexcept;
};

} // namespace sycl
Table 36. Constructors of the deprecated host buffer accessor
Constructor Description
template <typename AllocatorT>
accessor(buffer<DataT, 1, AllocatorT>& bufferRef,
         const property_list& propList = {})

Available only when (Dimensions == 0).

Constructs an accessor for accessing the first element of a buffer immediately on the host. The optional property_list provides properties for the constructed accessor.

template <typename AllocatorT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs an accessor for accessing a buffer immediately on the host. The optional property_list provides properties for the constructed accessor.

template <typename AllocatorT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         range<Dimensions> accessRange, const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs an accessor that is a ranged accessor which accesses a buffer immediately on the host, where the range starts at the beginning of the buffer. The optional property_list provides properties for the constructed accessor.

Throws an exception with the errc::invalid error code if accessRange exceeds the range of bufferRef in any dimension.

template <typename AllocatorT>
accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
         range<Dimensions> accessRange, id<Dimensions> accessOffset,
         const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs an accessor that is a ranged accessor which accesses a buffer immediately on the host, where the range starts at an offset from the beginning of the buffer. The optional property_list provides properties for the constructed accessor.

Throws an exception with the errc::invalid error code if the sum of accessRange and accessOffset exceeds the range of bufferRef in any dimension.

Table 37. Member functions of the deprecated host buffer accessor
Member function Description
bool is_placeholder() const

Always returns false.

id<Dimensions> get_offset() const

Available only when (Dimensions > 0).

If this is a ranged accessor, returns the offset that was specified when the accessor was constructed, otherwise returns the default constructed id<Dimensions>{}.

std::add_pointer_t<value_type> get_pointer() const noexcept

Returns a pointer to the start of this accessor’s underlying buffer, even if this is a ranged accessor whose range does not start at the beginning of the buffer. The return value is unspecified if the accessor is empty.

4.7.6.9.4.7. Accessor specialization with target::local

The accessor class may be specialized with target target::local, which results in a local accessor that has the same semantics and restrictions as local_accessor.

This specialization of accessor is only available for access modes access_mode::read_write and access_mode::atomic.

A synopsis for this specialization of accessor is provided below. Since some of the class types and member functions have the same name and meaning as other accessors, the common types and functions are described in Section 4.7.6.9.4.8. The member types are listed in Table 40. The constructors are listed in Table 38, and the member functions are listed in Table 41 and Table 39.

The additional common special member functions and common member functions are listed in Section 4.5.2 in Table 7 and Table 8, respectively. Additionally, accessors of the same type must be equality comparable.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
namespace sycl {

template <typename DataT, int Dimensions, access_mode AccessMode,
          target AccessTarget, access::placeholder IsPlaceholder>
class accessor {
 public:
  using value_type = DataT;
  using reference = DataT&;
  using const_reference = const DataT&;

  /* Available only when: (Dimensions == 0) */
  accessor(handler& commandGroupHandlerRef, const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  accessor(range<Dimensions> allocationSize, handler& commandGroupHandlerRef,
           const property_list& propList = {});

  /* -- common interface members -- */

  std::size_t get_size() const;

  std::size_t get_count() const;

  /* Available only when: (Dimensions > 0) */
  range<Dimensions> get_range() const;

  /* Available only when: (AccessMode == access_mode::read_write && Dimensions
   * == 0) */
  operator reference() const;

  /* Available only when: (AccessMode == access_mode::read_write && Dimensions >
   * 0) */
  reference operator[](id<Dimensions> index) const;

  /* Available only when: (Dimensions > 1) */
  __unspecified__ operator[](std::size_t index) const;

  /* Available only when: (AccessMode == access_mode::read_write && Dimensions
   * == 1) */
  reference operator[](std::size_t index) const;

  /* Available only when: (AccessMode == access_mode::atomic && Dimensions == 0)
   */
  operator atomic<DataT, access::address_space::local_space>() const;

  /* Available only when: (AccessMode == access_mode::atomic && Dimensions > 0)
   */
  atomic<DataT, access::address_space::local_space>
  operator[](id<Dimensions> index) const;

  /* Available only when: (AccessMode == access_mode::atomic && Dimensions == 1)
   */
  atomic<DataT, access::address_space::local_space>
  operator[](std::size_t index) const;

  local_ptr<DataT> get_pointer() const noexcept;
};

} // namespace sycl
Table 38. Constructors of the deprecated local accessor
Constructor Description
accessor(handler& commandGroupHandlerRef, const property_list& propList = {})

Available only when (Dimensions == 0).

Constructs an accessor instance for accessing local memory of a single DataT element within a SYCL kernel function on the queue associated with commandGroupHandlerRef. The optional property_list provides properties for the constructed accessor.

accessor(range<Dimensions> allocationSize, handler& commandGroupHandlerRef,
         const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs an accessor instance for accessing local memory of an array of DataT elements within a SYCL kernel function on the queue associated with commandGroupHandlerRef. The number of elements in the array is defined by allocationSize. The optional property_list provides properties for the constructed accessor.

Table 39. Member functions of the deprecated local accessor
Member function Description
operator atomic<DataT, access::address_space::local_space>() const

Available only when (AccessMode == access_mode::atomic && Dimensions == 0).

Returns an instance of atomic of type DataT providing atomic access to the element stored within the work-group’s local memory allocation that this accessor is accessing.

Preconditions: Must be called within a command.

atomic<DataT, access::address_space::local_space>
operator[](id<Dimensions> index) const

Available only when (AccessMode == access_mode::atomic && Dimensions > 0).

Returns an instance of atomic of type DataT providing atomic access to the element stored within the work-group’s local memory allocation that this accessor is accessing, at the index specified by index.

Preconditions: Must be called within a command.

atomic<DataT, access::address_space::local_space> operator[](std::size_t index) const

Available only when (AccessMode == access_mode::atomic && Dimensions == 1).

Returns an instance of atomic of type DataT providing atomic access to the element stored within the work-group’s local memory allocation that this accessor is accessing, at the index specified by index.

Preconditions: Must be called within a command.

local_ptr<DataT> get_pointer() const noexcept

Returns a multi_ptr to the work-group’s local memory allocation that this accessor is accessing. The return value is unspecified if the accessor is empty.

Preconditions: Must be called within a command.

4.7.6.9.4.8. Common members for deprecated accessors

Specializations of the accessor class with target::constant_buffer, target::host_buffer and target::local have many member types and member functions with the same name and meaning. Table 40 describes these common types and Table 41 describes the common member functions.

Table 40. Common member types of the deprecated accessors
Member types Description
value_type

If (AccessMode == access_mode::read), equal to const DataT, otherwise equal to DataT.

reference

Equal to value_type&.

const_reference

Equal to const DataT&.

Table 41. Common member functions of the deprecated accessors
Member function Description
std::size_t get_size() const noexcept

Returns the size in bytes of the memory region this accessor may access.

When AccessTarget is target::constant_buffer or target::host_buffer, the returned value is the size of the elements in the underlying buffer, unless this is a ranged accessor in which case it is the size of the elements within the accessor’s range.

When AccessTarget is target::local, the returned value is the size in bytes of the accessor’s local memory allocation, per work-group.

std::size_t get_count() const noexcept

Returns the number of DataT elements of the memory region this accessor may access.

When AccessTarget is target::constant_buffer or target::host_buffer, the returned value is the number of elements in the underlying buffer, unless this is a ranged accessor in which case it is the number of elements within the accessor’s range.

When AccessTarget is target::local, the returned value is the number of elements in the accessor’s local memory allocation, per work-group.

range<Dimensions> get_range() const

Available only when (Dimensions > 0).

Returns a range object which represents the number of elements of DataT per dimension that this accessor may access.

When AccessTarget is target::constant_buffer or target::host_buffer, the returned value is the range of the underlying buffer, unless this is a ranged accessor in which case it is the range that was specified when the accessor was constructed.

When AccessTarget is target::local, the returned value is the range that was specified when the accessor was constructed.

operator reference() const

When AccessTarget is target::constant_buffer or target::host_buffer, available only when (Dimensions == 0).

When AccessTarget is target::local, available only when (AccessMode == access_mode::read_write && Dimensions == 0).

Returns a reference to the single element that is accessed by this accessor.

When AccessTarget is target::local or target::constant_buffer, this function may only be called from within a command.

reference operator[](id<Dimensions> index) const

When AccessTarget is target::constant_buffer or target::host_buffer, available only when (Dimensions > 0).

When AccessTarget is target::local, available only when (AccessMode == access_mode::read_write && Dimensions > 0).

Returns a reference to the element at the location specified by index. If this is a ranged accessor, the element is determined by adding index to the accessor’s offset.

When AccessTarget is target::local or target::constant_buffer, this function may only be called from within a command.

__unspecified__ operator[](std::size_t index) const

Available only when (Dimensions > 1).

Returns an instance of an undefined intermediate type representing this accessor, with the dimensionality Dimensions-1 and containing an implicit id with index Dimensions set to index. The intermediate type returned must provide all available subscript operators which take a std::size_t parameter defined by this accessor class that are appropriate for the type it represents (including this subscript operator).

If this is a ranged accessor, the implicit id in the returned instance also includes the accessor’s offset.

When AccessTarget is target::local or target::constant_buffer, this function may only be called from within a command.

reference operator[](std::size_t index) const

When AccessTarget is target::constant_buffer or target::host_buffer, available only when (Dimensions == 1).

When AccessTarget is target::local, available only when (AccessMode == access_mode::read_write && Dimensions == 1).

Returns a reference to the element at the location specified by index. If this is a ranged accessor, the element is determined by adding index to the accessor’s offset.

When AccessTarget is target::local or target::constant_buffer, this function may only be called from within a command.

4.7.6.9.4.9. Accessor specialization with access_mode::atomic

The accessor class may be specialized with target target::device and access mode access_mode::atomic. This specialization provides additional member functions beyond those that are provided for other target::device specializations as described in Table 42.

Table 42. Deprecated atomic member functions of the accessor class
Member function Description
operator atomic<DataT, access::address_space::global_space>() const

Available only when (AccessMode == access_mode::atomic && Dimensions == 0).

Returns an instance of atomic of type DataT providing atomic access to the single element that is accessed by this accessor.

atomic<DataT, access::address_space::global_space>
operator[](id<Dimensions> index) const

Available only when (AccessMode == access_mode::atomic && Dimensions > 0).

Returns an instance of atomic of type DataT providing atomic access to the element stored within the accessor’s buffer at the index specified by index.

If this is a ranged accessor, the returned atomic instance provides access to the buffer element whose location is determined by adding the accessor’s offset to index.

atomic<DataT, access::address_space::global_space>
operator[](std::size_t index) const

Available only when (AccessMode == access_mode::atomic && Dimensions == 1).

Returns an instance of atomic of type DataT providing atomic access to the element stored within the accessor’s buffer at the index specified by index.

If this is a ranged accessor, the returned atomic instance provides access to the buffer element whose location is determined by adding the accessor’s offset to index.

4.7.6.10. Buffer accessor for host code

The host_accessor class provides access to data in a buffer from host code that is outside of a command (i.e. do not use this class to access a buffer inside a host task).

As with accessor, the dimensionality of host_accessor must match the underlying buffer, however, there is a special case if the buffer is one-dimensional. In this case, the accessor may either be one-dimensional or it may be zero-dimensional. A zero-dimensional accessor has access to just the first element of the buffer, whereas a one-dimensional accessor has access to the entire buffer.

The host_accessor class supports the following access modes: access_mode::read, access_mode::write and access_mode::read_write.

4.7.6.10.1. Interface for buffer host accessors

A synopsis of the host_accessor class is provided below. Since some of the class types and member functions have the same name and meaning as other accessors, the common types and functions are described in Section 4.7.6.12. The member types are listed in Table 51. The constructors are listed in Table 43, and the member functions are listed in Table 52 and Table 44.

The additional common special member functions and common member functions are listed in Section 4.5.2 in Table 7 and Table 8, respectively. For valid implicit conversions between accessor types refer to Section 4.7.6.10.3. Additionally, accessors of the same type must be equality comparable.

  1
  2
  3
  4
  5
  6
  7
  8
  9
 10
 11
 12
 13
 14
 15
 16
 17
 18
 19
 20
 21
 22
 23
 24
 25
 26
 27
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
namespace sycl {
template <typename DataT, int Dimensions = 1,
          access_mode AccessMode =
              (std::is_const_v<DataT> ? access_mode::read
                                      : access_mode::read_write)>
class host_accessor {
 public:
  using value_type = // const DataT for read-only accessors, DataT otherwise
      __value_type__;
  using reference = value_type&;
  using const_reference = const DataT&;
  using iterator = __unspecified_iterator__<value_type>;
  using const_iterator = __unspecified_iterator__<const value_type>;
  using reverse_iterator = std::reverse_iterator<iterator>;
  using const_reverse_iterator = std::reverse_iterator<const_iterator>;
  using difference_type =
      typename std::iterator_traits<iterator>::difference_type;
  using size_type = std::size_t;

  host_accessor();

  /* Available only when: (Dimensions == 0) */
  template <typename AllocatorT>
  host_accessor(buffer<DataT, 1, AllocatorT>& bufferRef,
                const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT>
  host_accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
                const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT, typename DeductionTagT>
  host_accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef, DeductionTagT tag,
                const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT>
  host_accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
                range<Dimensions> accessRange,
                const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT, typename DeductionTagT>
  host_accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
                range<Dimensions> accessRange, DeductionTagT tag,
                const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT>
  host_accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
                range<Dimensions> accessRange, id<Dimensions> accessOffset,
                const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  template <typename AllocatorT, typename DeductionTagT>
  host_accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
                range<Dimensions> accessRange, id<Dimensions> accessOffset,
                DeductionTagT tag, const property_list& propList = {});

  /* -- common interface members -- */

  void swap(host_accessor& other);

  size_type byte_size() const noexcept;

  size_type size() const noexcept;

  // Deprecated
  size_type max_size() const noexcept;

  bool empty() const noexcept;

  /* Available only when: (Dimensions > 0) */
  range<Dimensions> get_range() const;

  /* Available only when: (Dimensions > 0) */
  id<Dimensions> get_offset() const;

  /* Available only when: (Dimensions == 0) */
  operator reference() const;

  /* Available only when: (AccessMode != access_mode::read && Dimensions == 0) */
  const host_accessor& operator=(const value_type& other) const;

  /* Available only when: (AccessMode != access_mode::read && Dimensions == 0) */
  const host_accessor& operator=(value_type&& other) const;

  /* Available only when: (Dimensions > 0) */
  reference operator[](id<Dimensions> index) const;

  /* Available only when: (Dimensions > 1) */
  __unspecified__ operator[](std::size_t index) const;

  /* Available only when: (Dimensions == 1) */
  reference operator[](std::size_t index) const;

  std::add_pointer_t<value_type> get_pointer() const noexcept;

  iterator begin() const noexcept;

  iterator end() const noexcept;

  const_iterator cbegin() const noexcept;

  const_iterator cend() const noexcept;

  reverse_iterator rbegin() const noexcept;

  reverse_iterator rend() const noexcept;

  const_reverse_iterator crbegin() const noexcept;

  const_reverse_iterator crend() const noexcept;
};
} // namespace sycl
Table 43. Constructors of the host_accessor class
Constructor Description
host_accessor()

Constructs an empty accessor which fulfills the following post-conditions:

  • (empty() == true)

  • All size queries return 0.

  • The return value of get_pointer() is unspecified.

  • Trying to access the underlying memory is undefined behavior.

template <typename AllocatorT>
host_accessor(buffer<DataT, 1, AllocatorT>& bufferRef,
              const property_list& propList = {})

Available only when (Dimensions == 0).

Constructs a host_accessor for accessing the first element of a buffer immediately on the host. The optional property_list provides properties for the constructed accessor.

template <typename AllocatorT>
host_accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
              const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs a host_accessor for accessing a buffer immediately on the host. The optional property_list provides properties for the constructed accessor.

template <typename AllocatorT, typename DeductionTagT>
host_accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef, DeductionTagT tag,
              const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs a host_accessor for accessing a buffer immediately on the host. The tag is used to deduce template arguments of the accessor as described in Section 4.7.6.10.2. The optional property_list provides properties for the constructed accessor.

template <typename AllocatorT>
host_accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
              range<Dimensions> accessRange, const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs a host_accessor that is a ranged accessor which accesses a buffer immediately on the host, where the range starts at the beginning of the buffer. The optional property_list provides properties for the constructed accessor.

Throws an exception with the errc::invalid error code if accessRange exceeds the range of bufferRef in any dimension.

template <typename AllocatorT, typename DeductionTagT>
host_accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
              range<Dimensions> accessRange, DeductionTagT tag,
              const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs a host_accessor that is a ranged accessor which accesses a buffer immediately on the host, where the range starts at the beginning of the buffer. The tag is used to deduce template arguments of the accessor as described in Section 4.7.6.10.2. The optional property_list provides properties for the constructed accessor.

Throws an exception with the errc::invalid error code if accessRange exceeds the range of bufferRef in any dimension.

template <typename AllocatorT>
host_accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
              range<Dimensions> accessRange, id<Dimensions> accessOffset,
              const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs a host_accessor that is a ranged accessor which accesses a buffer immediately on the host, where the range starts at an offset from the beginning of the buffer. The optional property_list provides properties for the constructed accessor.

Throws an exception with the errc::invalid error code if the sum of accessRange and accessOffset exceeds the range of bufferRef in any dimension.

template <typename AllocatorT, typename DeductionTagT>
host_accessor(buffer<DataT, Dimensions, AllocatorT>& bufferRef,
              range<Dimensions> accessRange, id<Dimensions> accessOffset,
              DeductionTagT tag, const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs a host_accessor that is a ranged accessor which accesses a buffer immediately on the host, where the range starts at an offset from the beginning of the buffer. The tag is used to deduce template arguments of the accessor as described in Section 4.7.6.10.2. The optional property_list provides properties for the constructed accessor.

Throws an exception with the errc::invalid error code if the sum of accessRange and accessOffset exceeds the range of bufferRef in any dimension.

Table 44. Member functions of the host_accessor class
Member function Description
void swap(host_accessor& other);

Swaps the contents of the current accessor with the contents of other.

id<Dimensions> get_offset() const

Available only when (Dimensions > 0).

If this is a ranged accessor, returns the offset that was specified when the accessor was constructed. For other accessors, returns the default constructed id<Dimensions>{}.

std::add_pointer_t<value_type> get_pointer() const noexcept

Returns a pointer to the start of this accessor’s underlying buffer, even if this is a ranged accessor whose range does not start at the beginning of the buffer. The return value is unspecified if the accessor is empty.

const host_accessor& operator=(const value_type& other) const

Available only when (AccessMode != access_mode::read && Dimensions == 0).

Assignment to the single element that is accessed by this accessor.

const host_accessor& operator=(value_type&& other) const

Available only when (AccessMode != access_mode::read && Dimensions == 0).

Assignment to the single element that is accessed by this accessor.

4.7.6.10.2. Deduction tags for buffer host accessors

Some host_accessor constructors take a DeductionTagT parameter, which is used to deduce template arguments. The permissible values for this parameter are listed in Table 45 along with the access mode that they imply.

Table 45. Enumeration of tags available for host_accessor construction
Tag value Access mode

read_write

access_mode::read_write

read_only

access_mode::read

write_only

access_mode::write

4.7.6.10.3. Read only buffer host accessors and implicit conversions

Table 46 shows the specializations of host_accessor that are read-only accessors. There is an implicit conversion between any of these specializations, provided that all other template parameters are the same.

Table 46. Specializations of host_accessor that are read-only
Data type Access mode

not const-qualified

access_mode::read

const-qualified

access_mode::read

There is also an implicit conversion from the read-write host_accessor type shown in Table 47 to any of the read-only accessors in Table 46, provided that all other template parameters are the same.

Table 47. Specializations of host_accessor that are read-write
Data type Access mode

not const-qualified

access_mode::read_write

4.7.6.11. Local accessor

The local_accessor class allocates device local memory and provides access to this memory from within a SYCL kernel function. The local memory that is allocated is shared between all work-items of a work-group. If multiple work-groups execute simultaneously in an implementation, each work-group receives its own independent copy of the allocated local memory.

The underlying DataT type can be any C++ type that the device supports. If DataT is an implicit-lifetime type (as defined in the C++ core language), the local accessor implicitly creates objects of that type with indeterminate values. For other types, the local accessor merely allocates uninitialized memory, and the application is responsible for constructing objects in that memory (e.g. by calling placement-new).

A local accessor must not be used in a SYCL kernel function that is invoked via single_task or via the simple form of parallel_for that takes a range parameter. In these cases submitting the kernel to a queue must throw a synchronous exception with the errc::kernel_argument error code.

4.7.6.11.1. Interface for local accessors

A synopsis of the local_accessor class is provided below. Since some of the class types and member functions have the same name and meaning as other accessors, the common types and functions are described in Section 4.7.6.12. The member types are listed in Table 51 and Table 48. The constructors are listed in Table 49, and the member functions are listed in Table 52 and Table 50.

The additional common special member functions and common member functions are listed in Section 4.5.2 in Table 7 and Table 8, respectively. For valid implicit conversions between accessor types refer to Section 4.7.6.11.2. Additionally, accessors of the same type must be equality comparable.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
namespace sycl {
template <typename DataT, int Dimensions = 1> class local_accessor {
 public:
  using value_type = DataT;
  using reference = value_type&;
  using const_reference = const DataT&;
  template <access::decorated IsDecorated>
  using accessor_ptr =
      multi_ptr<value_type, access::address_space::local_space, IsDecorated>;
  using iterator = __unspecified_iterator__<value_type>;
  using const_iterator = __unspecified_iterator__<const value_type>;
  using reverse_iterator = std::reverse_iterator<iterator>;
  using const_reverse_iterator = std::reverse_iterator<const_iterator>;
  using difference_type =
      typename std::iterator_traits<iterator>::difference_type;
  using size_type = std::size_t;

  local_accessor();

  /* Available only when: (Dimensions == 0) */
  local_accessor(handler& commandGroupHandlerRef,
                 const property_list& propList = {});

  /* Available only when: (Dimensions > 0) */
  local_accessor(range<Dimensions> allocationSize,
                 handler& commandGroupHandlerRef,
                 const property_list& propList = {});

  /* -- common interface members -- */

  void swap(local_accessor& other);

  size_type byte_size() const noexcept;

  size_type size() const noexcept;

  // Deprecated
  size_type max_size() const noexcept;

  bool empty() const noexcept;

  range<Dimensions> get_range() const;

  /* Available only when: (Dimensions == 0) */
  operator reference() const;

  /* Available only when: (!std::is_const_v<DataT> && Dimensions == 0) */
  const local_accessor& operator=(const value_type& other) const;

  /* Available only when: (!std::is_const_v<DataT> && Dimensions == 0) */
  const local_accessor& operator=(value_type&& other) const;

  /* Available only when: (Dimensions > 0) */
  reference operator[](id<Dimensions> index) const;

  /* Available only when: (Dimensions > 1) */
  __unspecified__ operator[](std::size_t index) const;

  /* Available only when: (Dimensions == 1) */
  reference operator[](std::size_t index) const;

  /* Deprecated in SYCL 2020 */
  local_ptr<value_type> get_pointer() const noexcept;

  template <access::decorated IsDecorated>
  accessor_ptr<IsDecorated> get_multi_ptr() const noexcept;

  iterator begin() const noexcept;

  iterator end() const noexcept;

  const_iterator cbegin() const noexcept;

  const_iterator cend() const noexcept;

  reverse_iterator rbegin() const noexcept;

  reverse_iterator rend() const noexcept;

  const_reverse_iterator crbegin() const noexcept;

  const_reverse_iterator crend() const noexcept;
};
} // namespace sycl
Table 48. Member types of the local_accessor class
Member types Description
template <access::decorated IsDecorated> accessor_ptr

Equal to multi_ptr<value_type, access::address_space::local_space, IsDecorated>.

Table 49. Constructors of the local_accessor class
Constructor Description
local_accessor()

Constructs an empty local accessor which fulfills the following post-conditions:

  • (empty() == true)

  • All size queries return 0.

  • The return values of get_pointer() and get_multi_ptr() are unspecified.

  • A default constructed local accessor can be passed to a SYCL kernel function, but attempting to access data elements from it produces undefined behavior.

local_accessor(handler& commandGroupHandlerRef,
               const property_list& propList = {})

Available only when (Dimensions == 0).

Constructs a local_accessor for accessing local memory of a single DataT element within a SYCL kernel function on the queue associated with commandGroupHandlerRef. The optional property_list provides properties for the constructed accessor.

local_accessor(range<Dimensions> allocationSize,
               handler& commandGroupHandlerRef,
               const property_list& propList = {})

Available only when (Dimensions > 0).

Constructs a local_accessor for accessing local memory of an array of DataT elements within a SYCL kernel function on the queue associated with commandGroupHandlerRef. The number of elements in the array is defined by allocationSize. The optional property_list provides properties for the constructed accessor.

Table 50. Member functions of the local_accessor class
Member function Description
void swap(local_accessor& other);

Swaps the contents of the current accessor with the contents of other.

local_ptr<value_type> get_pointer() const noexcept

Returns a multi_ptr to the start of this accessor’s local memory region which corresponds to the calling work-group. The return value is unspecified if the accessor is empty.

Preconditions: Must be called within a command.

Deprecated in SYCL 2020. Use get_multi_ptr instead.

template <access::decorated IsDecorated>
accessor_ptr<IsDecorated> get_multi_ptr() const noexcept

Returns a multi_ptr to the start of the accessor’s local memory region which corresponds to the calling work-group. The return value is unspecified if the accessor is empty.

This function may only be called from within a SYCL kernel function.

const local_accessor& operator=(const value_type& other) const

Available only when (!std::is_const_v<DataT> && Dimensions == 0).

Assignment to the single element that is accessed by this accessor.

Preconditions: Must be called within a command.

const local_accessor& operator=(const value_type&& other) const

Available only when (!std::is_const_v<DataT> && Dimensions == 0).

Assignment to the single element that is accessed by this accessor.

Preconditions: Must be called within a command.

4.7.6.11.2. Read only local accessors and implicit conversions

Since local_accessor has no template parameter for the access mode, the only specialization for a read-only local accessor is by providing a const qualified DataT parameter. Specializations with a non-const qualified DataT parameter are read-write. There is an implicit conversion from the read-write specialization to the read-only specialization, provided that all other template parameters are the same.

4.7.6.12. Common members for buffer and local accessors

The accessor, host_accessor, and local_accessor classes have many member types and member functions with the same name and meaning. Table 51 describes these common types and Table 52 describes the common member functions.

Table 51. Common buffer and local accessor member types
Member types Description
value_type

If the accessor is read-only, equal to const DataT, otherwise equal to DataT.

See Section 4.7.6.9.3, Section 4.7.6.10.3 and Section 4.7.6.11.2 for which accessors are considered read-only.

reference

Equal to value_type&.

const_reference

Equal to const DataT&.

iterator

Iterator that can provide ranged access. Cannot be written to if the accessor is read-only. The underlying pointer is address space qualified for accessor specializations with target::device and for local_accessor.

const_iterator

Iterator that can provide ranged access. Cannot be written to. The underlying pointer is address space qualified for accessor specializations with target::device and for local_accessor.

reverse_iterator

Iterator adaptor that reverses the direction of iterator.

const_reverse_iterator

Iterator adaptor that reverses the direction of const_iterator.

difference_type

Equal to typename std::iterator_traits<iterator>::difference_type.

size_type

Equal to std::size_t.

Table 52. Common buffer and local accessor member functions
Member function Description
size_type byte_size() const noexcept

Returns the size in bytes of the memory region this accessor may access.

For a buffer accessor this is the size of the underlying buffer, unless it is a ranged accessor in which case it is the size of the elements within the accessor’s range.

For a local accessor this is the size of the accessor’s local memory allocation, per work-group.

size_type size() const noexcept

Returns the number of DataT elements of the memory region this accessor may access.

For a buffer accessor this is the number of elements in the underlying buffer, unless it is a ranged accessor in which case it is the number of elements within the accessor’s range.

For a local accessor this is the number of elements in the accessor’s local memory allocation, per work-group.

size_type max_size() const noexcept

Deprecated by SYCL 2020.

Returns the maximum number of elements any accessor of this type would be able to access.

bool empty() const noexcept

Returns true if (size() == 0).

range<Dimensions> get_range() const

Available only when (Dimensions > 0).

Returns a range object which represents the number of elements of DataT per dimension that this accessor may access.

For a buffer accessor this is the range of the underlying buffer, unless it is a ranged accessor in which case it is the range that was specified when the accessor was constructed.

operator reference() const

For accessor available only when (AccessMode != access_mode::atomic && Dimensions == 0).

For host_accessor and local_accessor available only when (Dimensions == 0).

Returns a reference to the single element that is accessed by this accessor.

For accessor and local_accessor, this function may only be called from within a command.

reference operator[](id<Dimensions> index) const

For accessor available only when (AccessMode != access_mode::atomic && Dimensions > 0).

For host_accessor and local_accessor available only when (Dimensions > 0).

Returns a reference to the element at the location specified by index. If this is a ranged accessor, the element is determined by adding index to the accessor’s offset.

For accessor and local_accessor, this function may only be called from within a command.

__unspecified__ operator[](std::size_t index) const

Available only when (Dimensions > 1).

Returns an instance of an undefined intermediate type representing this accessor, with the dimensionality Dimensions-1 and containing an implicit id with index Dimensions set to index. The intermediate type returned must provide all available subscript operators which take a std::size_t parameter defined by this accessor class that are appropriate for the type it represents (including this subscript operator).

If this is a ranged accessor, the implicit id in the returned instance also includes the accessor’s offset.

For accessor and local_accessor, this function may only be called from within a command.

reference operator[](std::size_t index) const

For accessor available only when (AccessMode != access_mode::atomic && Dimensions == 1).

For host_accessor and local_accessor available only when (Dimensions == 1).

Returns a reference to the element at the location specified by index. If this is a ranged accessor, the element is determined by adding index to the accessor’s offset.

For accessor and local_accessor, this function may only be called from within a command.

iterator begin() const noexcept

Returns an iterator to the first element of the memory this accessor may access.

For a buffer accessor this is an iterator to the first element of the underlying buffer, unless this is a ranged accessor in which case it is an iterator to first element within the accessor’s range.

For accessor and local_accessor, this function may only be called from within a command.

iterator end() const noexcept

Returns an iterator to one element past the last element of the memory this accessor may access.

For a buffer accessor this is an iterator to one element past the last element in the underlying buffer, unless this is a ranged accessor in which case it is an iterator to one element past the last element within the accessor’s range.

For accessor and local_accessor, this function may only be called from within a command.

const_iterator cbegin() const noexcept

Returns a const iterator to the first element of the memory this accessor may access.

For a buffer accessor this is a const iterator to the first element of the underlying buffer, unless this is a ranged accessor in which case it is a const iterator to first element within the accessor’s range.

For accessor and local_accessor, this function may only be called from within a command.

const_iterator cend() const noexcept

Returns a const iterator to one element past the last element of the memory this accessor may access.

For a buffer accessor this is a const iterator to one element past the last element in the underlying buffer, unless this is a ranged accessor in which case it is a const iterator to one element past the last element within the accessor’s range.

For accessor and local_accessor, this function may only be called from within a command.

reverse_iterator rbegin() const noexcept

Returns an iterator adaptor to the last element of the memory this accessor may access.

For a buffer accessor this is an iterator adaptor to the last element of the underlying buffer, unless this is a ranged accessor in which case it is an iterator adaptor to the last element within the accessor’s range.

For accessor and local_accessor, this function may only be called from within a command.

reverse_iterator rend() const noexcept

Returns an iterator adaptor to one element before the first element of the memory this accessor may access.

For a buffer accessor this is an iterator adaptor to one element before the first element in the underlying buffer, unless this is a ranged accessor in which case it is an iterator adaptor to one element before the first element within the accessor’s range.

For accessor and local_accessor, this function may only be called from within a command.

const_reverse_iterator crbegin() const noexcept

Returns a const iterator adaptor to the last element of the memory this accessor may access.

For a buffer accessor this is a const iterator adaptor to the last element of the underlying buffer, unless this is a ranged accessor in which case it is an const iterator adaptor to last element within the accessor’s range.

For accessor and local_accessor, this function may only be called from within a command.

const_reverse_iterator crend() const noexcept

Returns a const iterator adaptor to one element before the first element of the memory this accessor may access.

For a buffer accessor this is a const iterator adaptor to one element before the first element in the underlying buffer, unless this is a ranged accessor in which case it is a const iterator adaptor to one element before the first element within the accessor’s range.

For accessor and local_accessor, this function may only be called from within a command.

4.7.6.13. Unsampled image accessors

There are two classes which implement accessors for unsampled images, unsampled_image_accessor and host_unsampled_image_accessor. The former provides access from within a SYCL kernel function or from within a host task. The latter provides access from host code that is outside of a host task.

The dimensionality of an unsampled image accessor must match the dimensionality of the underlying image to which it provides access. Both unsampled image accessor classes support the access_mode::read and access_mode::write access modes. In addition, the host_unsampled_image_accessor class supports access_mode::read_write.

The AccessTarget template parameter dictates how the unsampled_image_accessor can be used: image_target::device means the accessor can be used in a SYCL kernel function while image_target::host_task means the accessor can be used in a host task. Programs which specify this template parameter as image_target::device and then use the unsampled_image_accessor from a host task are ill formed. Likewise, programs which specify this template parameter as image_target::host_task and then use the unsampled_image_accessor from a SYCL kernel function are ill formed.

4.7.6.13.1. Interface for unsampled image accessors

A synopsis of the two unsampled image accessor classes is provided below. Both classes have member types with the same name, which are described in Table 53. The constructors for the two classes are described in Table 54 and Table 55. Both classes also have member functions with the same name, which are described in Table 56.

The additional common special member functions and common member functions are listed in Section 4.5.2 in Table 7 and Table 8, respectively. For valid implicit conversions between unsampled accessor types refer to Section 4.7.6.13.2.

Two unsampled_image_accessor objects of the same type must be equality comparable in both the host code and in SYCL kernel functions. Two host_unsampled_image_accessor objects of the same type must be equality comparable in the host code.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
namespace sycl {

enum class image_target : /* unspecified */ { device, host_task };

template <typename DataT, int Dimensions, access_mode AccessMode,
          image_target AccessTarget = image_target::device>
class unsampled_image_accessor {
 public:
  using value_type = // const DataT for read-only accessors, DataT otherwise
      __value_type__;
  using reference = value_type&;
  using const_reference = const DataT&;

  template <typename AllocatorT>
  unsampled_image_accessor(unsampled_image<Dimensions, AllocatorT>& imageRef,
                           handler& commandGroupHandlerRef,
                           const property_list& propList = {});

  /* -- common interface members -- */

  /* -- property interface members -- */

  std::size_t size() const noexcept;

  /* Available only when: AccessMode == access_mode::read
  if Dimensions == 1, CoordT = int
  if Dimensions == 2, CoordT = int2
  if Dimensions == 3, CoordT = int4 */
  template <typename CoordT> DataT read(const CoordT& coords) const noexcept;

  /* Available only when: AccessMode == access_mode::write
  if Dimensions == 1, CoordT = int
  if Dimensions == 2, CoordT = int2
  if Dimensions == 3, CoordT = int4 */
  template <typename CoordT>
  void write(const CoordT& coords, const DataT& color) const;
};

template <typename DataT, int Dimensions = 1,
          access_mode AccessMode =
              (std::is_const_v<DataT> ? access_mode::read
                                      : access_mode::read_write)>
class host_unsampled_image_accessor {
 public:
  using value_type = // const DataT for read-only accessors, DataT otherwise
      __value_type__;
  using reference = value_type&;
  using const_reference = const DataT&;

  template <typename AllocatorT>
  host_unsampled_image_accessor(
      unsampled_image<Dimensions, AllocatorT>& imageRef,
      const property_list& propList = {});

  /* -- common interface members -- */

  /* -- property interface members -- */

  std::size_t size() const noexcept;

  /* Available only when: (AccessMode == access_mode::read ||
                           AccessMode == access_mode::read_write)
  if Dimensions == 1, CoordT = int
  if Dimensions == 2, CoordT = int2
  if Dimensions == 3, CoordT = int4 */
  template <typename CoordT> DataT read(const CoordT& coords) const noexcept;

  /* Available only when: (AccessMode == access_mode::write ||
                           AccessMode == access_mode::read_write)
  if Dimensions == 1, CoordT = int
  if Dimensions == 2, CoordT = int2
  if Dimensions == 3, CoordT = int4 */
  template <typename CoordT>
  void write(const CoordT& coords, const DataT& color) const;
};

} // namespace sycl
Table 53. Member types of the unsampled image classes
Member types Description
value_type

If the accessor is read-only, equal to const DataT, otherwise equal to DataT.

See Section 4.7.6.13.2 for which accessors are considered read-only.

reference

Equal to value_type&.

const_reference

Equal to const DataT&.

Table 54. Constructors of the unsampled_image_accessor class
Constructor Description
template <typename AllocatorT>
unsampled_image_accessor(unsampled_image<Dimensions, AllocatorT>& imageRef,
                         handler& commandGroupHandlerRef,
                         const property_list& propList = {})

Constructs an unsampled_image_accessor for accessing an unsampled_image within a command on the queue associated with commandGroupHandlerRef. The optional property_list provides properties for the constructed object.

If AccessTarget is image_target::device, throws an exception with the errc::feature_not_supported error code if the device associated with commandGroupHandlerRef does not have aspect::image.

Table 55. Constructors of the host_unsampled_image_accessor class
Constructor Description
template <typename AllocatorT>
host_unsampled_image_accessor(unsampled_image<Dimensions, AllocatorT>& imageRef,
                              const property_list& propList = {})

Constructs a host_unsampled_image_accessor for accessing an unsampled_image immediately on the host. The optional property_list provides properties for the constructed object.

Table 56. Member functions of the unsampled image classes
Member function Description
std::size_t size() const noexcept

Returns the number of elements of the underlying unsampled_image that this accessor is accessing.

template <typename CoordT> DataT read(const CoordT& coords) const

Available only when (AccessMode == access_mode::read || AccessMode == access_mode::read_write).

Reads and returns an element of the unsampled_image at the coordinates specified by coords. Permitted types for CoordT are int when Dimensions == 1, int2 when Dimensions == 2 and int4 when Dimensions == 3.

For unsampled_image_accessor, this function may only be called from within a command.

template <typename CoordT>
void write(const CoordT& coords, const DataT& color) const

Available only when (AccessMode == access_mode::write || AccessMode == access_mode::read_write).

Writes the value specified by color to the element of the image at the coordinates specified by coords. Permitted types for CoordT are int when Dimensions == 1, int2 when Dimensions == 2 and int4 when Dimensions == 3.

For unsampled_image_accessor, this function may only be called from within a command.

4.7.6.13.2. Read only unsampled image accessors and implicit conversions

All specializations of unsampled image accessors with access_mode::read are read-only regardless of whether DataT is const qualified. There is an implicit conversion between the const qualified and non-const qualified specializations, provided that all other template parameters are the same.

4.7.6.14. Sampled image accessors

There are two classes which implement accessors for sampled images, sampled_image_accessor and host_sampled_image_accessor. The former provides access from within a SYCL kernel function or from within a host task. The latter provides access from host code that is outside of a host task.

The dimensionality of a sampled image accessor must match the dimensionality of the underlying image to which it provides access. Sampled image accessors are always read-only.

The AccessTarget template parameter dictates how the sampled_image_accessor can be used: image_target::device means the accessor can be used in a SYCL kernel function while image_target::host_task means the accessor can be used in a host task. Programs which specify this template parameter as image_target::device and then use the sampled_image_accessor from a host task are ill formed. Likewise, programs which specify this template parameter as image_target::host_task and then use the sampled_image_accessor from a SYCL kernel function are ill formed.

4.7.6.14.1. Interface for sampled image accessors

A synopsis of the two sampled image accessor classes is provided below. Both classes have member types with the same name, which are described in Table 57. The constructors for the two classes are described in Table 58 and Table 59. Both classes also have member functions with the same name, which are described in Table 60.

The additional common special member functions and common member functions are listed in Section 4.5.2 in Table 7 and Table 8, respectively. For valid implicit conversions between sampled accessor types refer to Section 4.7.6.14.2.

Two sampled_image_accessor objects of the same type must be equality comparable in both the host code and in SYCL kernel functions. Two host_sampled_image_accessor objects of the same type must be equality comparable in the host code.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
namespace sycl {

enum class image_target : /* unspecified */ { device, host_task };

template <typename DataT, int Dimensions,
          image_target AccessTarget = image_target::device>
class sampled_image_accessor {
 public:
  using value_type = const DataT;
  using reference = const DataT&;
  using const_reference = const DataT&;

  template <typename AllocatorT>
  sampled_image_accessor(sampled_image<Dimensions, AllocatorT>& imageRef,
                         handler& commandGroupHandlerRef,
                         const property_list& propList = {});


  /* -- common interface members -- */

  /* -- property interface members -- */

  std::size_t size() const noexcept;

  /* if Dimensions == 1, CoordT = float
     if Dimensions == 2, CoordT = float2
     if Dimensions == 3, CoordT = float4 */
  template <typename CoordT> DataT read(const CoordT& coords) const noexcept;
};

template <typename DataT, int Dimensions> class host_sampled_image_accessor {
 public:
  using value_type = const DataT;
  using reference = const DataT&;
  using const_reference = const DataT&;

  template <typename AllocatorT>
  host_sampled_image_accessor(sampled_image<Dimensions, AllocatorT>& imageRef,
                              const property_list& propList = {});

  /* -- common interface members -- */

  /* -- property interface members -- */

  std::size_t size() const noexcept;

  /* if Dimensions == 1, CoordT = float
     if Dimensions == 2, CoordT = float2
     if Dimensions == 3, CoordT = float4 */
  template <typename CoordT> DataT read(const CoordT& coords) const noexcept;
};

} // namespace sycl
Table 57. Member types of the sampled image classes
Member types Description
value_type

Equal to const DataT.

reference

Equal to const DataT&.

const_reference

Equal to const DataT&.

Table 58. Constructors of the sampled_image_accessor class
Constructor Description
template <typename AllocatorT>
sampled_image_accessor(sampled_image<Dimensions, AllocatorT>& imageRef,
                       handler& commandGroupHandlerRef,
                       const property_list& propList = {})

Constructs a sampled_image_accessor for accessing a sampled_image within a command on the queue associated with commandGroupHandlerRef. The optional property_list provides properties for the constructed object.

If AccessTarget is image_target::device, throws an exception with the errc::feature_not_supported error code if the device associated with commandGroupHandlerRef does not have aspect::image.

Table 59. Constructors of the host_sampled_image_accessor class
Constructor Description
template <typename AllocatorT>
host_sampled_image_accessor(sampled_image<Dimensions, AllocatorT>& imageRef,
                            const property_list& propList = {})

Constructs a host_sampled_image_accessor for accessing a sampled_image immediately on the host. The optional property_list provides properties for the constructed object.

Table 60. Member functions of the sampled image classes
Member function Description
std::size_t size() const noexcept

Returns the number of elements of the underlying sampled_image that this accessor is accessing.

template <typename CoordT> DataT read(const CoordT& coords) const

Reads and returns a sampled element of the sampled_image at the coordinates specified by coords. Permitted types for CoordT are float when Dimensions == 1, float2 when Dimensions == 2 and float4 when Dimensions == 3.

For sampled_image_accessor, this function may only be called from within a command.

4.7.6.14.2. Read only sampled image accessors and implicit conversions

All specializations of sampled image accessors are read-only regardless of whether DataT is const qualified. There is an implicit conversion between the const qualified and non-const qualified specializations, provided that all other template parameters are the same.

4.7.7. Address space classes

In SYCL, there are five different address spaces: global, local, constant, private and generic. In a SYCL generic implementation, types are not affected by the address spaces. However, there are situations where users need to explicitly carry address spaces in the type. For example:

  • For performance tuning and genericness. Even if the platform supports the representation of the generic address space, this may come at some performance sacrifice. In order to help the target compiler, it can be useful to track specifically which address space a pointer is addressing.

  • When linking SYCL kernels with SYCL backend-specific functions. In this case, it might be necessary to specify the address space for any pointer parameters.

Direct declaration of pointers with address spaces is discouraged as the definition is implementation-defined. Users must rely on the multi_ptr class to handle address space boundaries and interoperability.

4.7.7.1. Multi-pointer class

The multi-pointer class is the common interface for the explicit pointer classes, defined in Section 4.7.7.2.

There are situations where a user may want to make their type address space dependent. This allows performing generic programming that depends on the address space associated with their data. An example might be wrapping a pointer inside a class, where a user may need to template the class according to the address space of the pointer the class is initialized with. In this case, the multi_ptr class enables users to do this in a portable and stable way.

The multi_ptr class exposes 3 flavors of the same interface. If the value of access::decorated is access::decorated::no, the interface exposes pointers and references type that are not decorated by an address space. If the value of access::decorated is access::decorated::yes, the interface exposes pointers and references type that are decorated by an address space. The decoration is implementation dependent and relies on device compiler extensions. The decorated type may be distinct from the non-decorated one. For interoperability with the SYCL backend, users should rely on types exposed by the decorated version. If the value of access::decorated is access::decorated::legacy, the 1.2.1 interface is exposed.

The template traits remove_decoration and type alias remove_decoration_t retrieve the non-decorated pointer or reference from a decorated one. Using this template trait with a non-decorated type is safe and returns the same type.

It is possible to use the void type for the multi_ptr class, but in that case some functionality is disabled. multi_ptr<void> does not provide the reference or const_reference types, the access operators (operator*(), operator->()), the arithmetic operators or prefetch member function. Conversions from multi_ptr to multi_ptr<void> of the same address space are allowed, and will occur implicitly. Conversions from multi_ptr<void> to any other multi_ptr type of the same address space are allowed, but must be explicit. The same rules apply to multi_ptr<const void>.

An overview of the interface provided for the multi_ptr class follows.

  1
  2
  3
  4
  5
  6
  7
  8
  9
 10
 11
 12
 13
 14
 15
 16
 17
 18
 19
 20
 21
 22
 23
 24
 25
 26
 27
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
namespace sycl {
namespace access {

enum class address_space : /* unspecified */ {
  global_space,
  local_space,
  constant_space, // Deprecated in SYCL 2020
  private_space,
  generic_space
};

enum class decorated : /* unspecified */ {
  no,
  yes,
  legacy
};

} // namespace access

template <typename T> struct remove_decoration {
  using type = /* ... */;
};

template <typename T> using remove_decoration_t = remove_decoration<T>::type;

template <typename ElementType, access::address_space Space,
          access::decorated DecorateAddress = access::decorated::legacy>
class multi_ptr {
 public:
  static constexpr bool is_decorated =
      DecorateAddress == access::decorated::yes;
  static constexpr access::address_space address_space = Space;

  using value_type = ElementType;
  using pointer = std::conditional_t<is_decorated, __unspecified__*,
                                     std::add_pointer_t<value_type>>;
  using reference = std::conditional_t<is_decorated, __unspecified__&,
                                       std::add_lvalue_reference_t<value_type>>;
  using iterator_category = std::random_access_iterator_tag;
  using difference_type = std::ptrdiff_t;

  static_assert(std::is_same_v<remove_decoration_t<pointer>,
                               std::add_pointer_t<value_type>>);
  static_assert(std::is_same_v<remove_decoration_t<reference>,
                               std::add_lvalue_reference_t<value_type>>);
  // Legacy has a different interface.
  static_assert(DecorateAddress != access::decorated::legacy);

  // Constructors
  multi_ptr();
  multi_ptr(const multi_ptr&);
  multi_ptr(multi_ptr&&);
  explicit multi_ptr(
      typename multi_ptr<ElementType, Space, access::decorated::yes>::pointer);
  multi_ptr(std::nullptr_t);

  // Available only when:
  //   (Space == access::address_space::global_space ||
  //    Space == access::address_space::generic_space) &&
  //   (std::is_same_v<std::remove_const_t<ElementType>, std::remove_const_t<AccDataT>>) &&
  //   (std::is_const_v<ElementType> ||
  //    !std::is_const_v<accessor<AccDataT, Dimensions, Mode, target::device,
  //                              IsPlaceholder>::value_type>)
  template <typename AccDataT, int Dimensions, access_mode Mode,
            access::placeholder IsPlaceholder>
  multi_ptr(
      accessor<AccDataT, Dimensions, Mode, target::device, IsPlaceholder>);

  // Available only when:
  //   (Space == access::address_space::local_space ||
  //    Space == access::address_space::generic_space) &&
  //   (std::is_same_v<std::remove_const_t<ElementType>, std::remove_const_t<AccDataT>>) &&
  //   (std::is_const_v<ElementType> || !std::is_const_v<AccDataT>)
  template <typename AccDataT, int Dimensions>
  multi_ptr(local_accessor<AccDataT, Dimensions>);

  // Deprecated
  // Available only when:
  //   (Space == access::address_space::local_space ||
  //    Space == access::address_space::generic_space) &&
  //   (std::is_same_v<std::remove_const_t<ElementType>, std::remove_const_t<AccDataT>>) &&
  //   (std::is_const_v<ElementType> || !std::is_const_v<AccDataT>)
  template <typename AccDataT, int Dimensions, access_mode Mode,
            access::placeholder IsPlaceholder>
  multi_ptr(
      accessor<AccDataT, Dimensions, Mode, target::local, IsPlaceholder>);

  // Deprecated
  // Available only when:
  //   Space == access::address_space::constant_space &&
  //   (std::is_same_v<std::remove_const_t<ElementType>, std::remove_const_t<AccDataT>>) &&
  //   (std::is_const_v<ElementType> || !std::is_const_v<AccDataT>)
  template <typename AccDataT, int Dimensions, access::placeholder IsPlaceholder>
  multi_ptr(
      accessor<AccDataT, Dimensions, access_mode::read, target::constant_buffer, IsPlaceholder>);

  // Assignment and access operators
  multi_ptr& operator=(const multi_ptr&);
  multi_ptr& operator=(multi_ptr&&);
  multi_ptr& operator=(std::nullptr_t);

  // Available only when:
  //   (Space == access::address_space::generic_space &&
  //    AS != access::address_space::constant_space)
  template <access::address_space AS, access::decorated IsDecorated>
  multi_ptr& operator=(const multi_ptr<value_type, AS, IsDecorated>&);

  // Available only when:
  //   (Space == access::address_space::generic_space &&
  //    AS != access::address_space::constant_space)
  template <access::address_space AS, access::decorated IsDecorated>
  multi_ptr& operator=(multi_ptr<value_type, AS, IsDecorated>&&);

  reference operator[](std::ptrdiff_t) const;

  reference operator*() const;
  pointer operator->() const;

  pointer get() const;
  std::add_pointer_t<value_type> get_raw() const;
  __unspecified__* get_decorated() const;

  // Conversion to the underlying pointer type
  // Deprecated, get() should be used instead.
  operator pointer() const;

  // Cast to private_ptr
  // Available only when: (Space == access::address_space::generic_space)
  template <access::decorated IsDecorated>
  explicit operator multi_ptr<value_type, access::address_space::private_space,
                              IsDecorated>() const;

  // Cast to private_ptr of const data
  // Available only when: (Space == access::address_space::generic_space)
  template <access::decorated IsDecorated>
  explicit operator multi_ptr<const value_type, access::address_space::private_space,
                              IsDecorated>() const;

  // Cast to global_ptr
  // Available only when: (Space == access::address_space::generic_space)
  template <access::decorated IsDecorated>
  explicit operator multi_ptr<value_type, access::address_space::global_space,
                              IsDecorated>() const;

  // Cast to global_ptr of const data
  // Available only when: (Space == access::address_space::generic_space)
  template <access::decorated IsDecorated>
  explicit operator multi_ptr<const value_type, access::address_space::global_space,
                              IsDecorated>() const;

  // Cast to local_ptr
  // Available only when: (Space == access::address_space::generic_space)
  template <access::decorated IsDecorated>
  explicit operator multi_ptr<value_type, access::address_space::local_space,
                              IsDecorated>() const;

  // Cast to local_ptr of const data
  // Available only when: (Space == access::address_space::generic_space)
  template <access::decorated IsDecorated>
  explicit operator multi_ptr<const value_type, access::address_space::local_space,
                              IsDecorated>() const;

  // Implicit conversion to a multi_ptr<void>.
  // Available only when: (!std::is_const_v<value_type>)
  template <access::decorated IsDecorated>
  operator multi_ptr<void, Space, IsDecorated>() const;

  // Implicit conversion to a multi_ptr<const void>.
  // Available only when: (std::is_const_v<value_type>)
  template <access::decorated IsDecorated>
  operator multi_ptr<const void, Space, IsDecorated>() const;

  // Implicit conversion to multi_ptr<const value_type, Space>.
  template <access::decorated IsDecorated>
  operator multi_ptr<const value_type, Space, IsDecorated>() const;

  // Implicit conversion to the non-decorated version of multi_ptr.
  // Available only when: (is_decorated == true)
  operator multi_ptr<value_type, Space, access::decorated::no>() const;

  // Implicit conversion to the decorated version of multi_ptr.
  // Available only when: (is_decorated == false)
  operator multi_ptr<value_type, Space, access::decorated::yes>() const;

  // Available only when: (Space == address_space::global_space)
  void prefetch(std::size_t numElements) const;

  // Arithmetic operators
  friend multi_ptr& operator++(multi_ptr& mp) { /* ... */
  }
  friend multi_ptr operator++(multi_ptr& mp, int) { /* ... */
  }
  friend multi_ptr& operator--(multi_ptr& mp) { /* ... */
  }
  friend multi_ptr operator--(multi_ptr& mp, int) { /* ... */
  }
  friend multi_ptr& operator+=(multi_ptr& lhs, difference_type r) { /* ... */
  }
  friend multi_ptr& operator-=(multi_ptr& lhs, difference_type r) { /* ... */
  }
  friend multi_ptr operator+(const multi_ptr& lhs,
                             difference_type r) { /* ... */
  }
  friend multi_ptr operator-(const multi_ptr& lhs,
                             difference_type r) { /* ... */
  }
  friend reference operator*(const multi_ptr& lhs) { /* ... */
  }

  friend bool operator==(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator!=(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator<(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator>(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator<=(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator>=(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }

  friend bool operator==(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator!=(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator<(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator>(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator<=(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator>=(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }

  friend bool operator==(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator!=(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator<(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator>(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator<=(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator>=(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
};

// Specialization of multi_ptr for void and const void
// VoidType can be either void or const void
template <access::address_space Space, access::decorated DecorateAddress>
class multi_ptr<VoidType, Space, DecorateAddress> {
 public:
  static constexpr bool is_decorated =
      DecorateAddress == access::decorated::yes;
  static constexpr access::address_space address_space = Space;

  using value_type = VoidType;
  using pointer = std::conditional_t<is_decorated, __unspecified__*,
                                     std::add_pointer_t<value_type>>;
  using difference_type = std::ptrdiff_t;

  static_assert(std::is_same_v<remove_decoration_t<pointer>,
                               std::add_pointer_t<value_type>>);
  // Legacy has a different interface.
  static_assert(DecorateAddress != access::decorated::legacy);

  // Constructors
  multi_ptr();
  multi_ptr(const multi_ptr&);
  multi_ptr(multi_ptr&&);
  explicit multi_ptr(
      typename multi_ptr<VoidType, Space, access::decorated::yes>::pointer);
  multi_ptr(std::nullptr_t);

  // Available only when:
  //   (Space == access::address_space::global_space ||
  //    Space == access::address_space::generic_space) &&
  //   (std::is_const_v<VoidType> ||
  //    !std::is_const_v<accessor<ElementType, Dimensions, Mode, target::device,
  //                              IsPlaceholder>::value_type>)
  template <typename ElementType, int Dimensions, access_mode Mode,
            access::placeholder IsPlaceholder>
  multi_ptr(
      accessor<ElementType, Dimensions, Mode, target::device, IsPlaceholder>);

  // Available only when:
  //   (Space == access::address_space::local_space ||
  //    Space == access::address_space::generic_space) &&
  //   (std::is_const_v<VoidType> || !std::is_const_v<ElementType>)
  template <typename ElementType, int Dimensions>
  multi_ptr(local_accessor<ElementType, Dimensions>);

  // Deprecated
  // Available only when:
  //   (Space == access::address_space::local_space ||
  //    Space == access::address_space::generic_space) &&
  //   (std::is_const_v<VoidType> || !std::is_const_v<ElementType>)
  template <typename ElementType, int Dimensions, access_mode Mode,
            access::placeholder IsPlaceholder>
  multi_ptr(
      accessor<ElementType, Dimensions, Mode, target::local, IsPlaceholder>);

  // Deprecated
  // Available only when:
  //   Space == access::address_space::constant_space &&
  //   (std::is_const_v<VoidType> || !std::is_const_v<ElementType>)
  template <typename ElementType, int Dimensions, access::placeholder IsPlaceholder>
  multi_ptr(
      accessor<ElementType, Dimensions, access_mode::read, target::constant_buffer, IsPlaceholder>);

  // Assignment operators
  multi_ptr& operator=(const multi_ptr&);
  multi_ptr& operator=(multi_ptr&&);
  multi_ptr& operator=(std::nullptr_t);

  pointer get() const;

  // Conversion to the underlying pointer type
  operator pointer() const;

  // Explicit conversion to a multi_ptr<ElementType>
  // Available only when: (std::is_const_v<ElementType> || !std::is_const_v<VoidType>)
  template <typename ElementType>
  explicit operator multi_ptr<ElementType, Space, DecorateAddress>() const;

  // Implicit conversion to the non-decorated version of multi_ptr.
  // Available only when: (is_decorated == true)
  operator multi_ptr<value_type, Space, access::decorated::no>() const;

  // Implicit conversion to the decorated version of multi_ptr.
  // Available only when: (is_decorated == false)
  operator multi_ptr<value_type, Space, access::decorated::yes>() const;

  // Implicit conversion to multi_ptr<const void, Space>
  operator multi_ptr<const void, Space, DecorateAddress>() const;

  friend bool operator==(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator!=(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator<(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator>(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator<=(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator>=(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }

  friend bool operator==(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator!=(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator<(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator>(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator<=(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator>=(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }

  friend bool operator==(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator!=(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator<(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator>(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator<=(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator>=(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
};

// Deprecated, address_space_cast should be used instead.
template <typename ElementType, access::address_space Space,
          access::decorated DecorateAddress>
multi_ptr<ElementType, Space, DecorateAddress> make_ptr(ElementType*);

template <access::address_space Space, access::decorated DecorateAddress,
          typename ElementType>
multi_ptr<ElementType, Space, DecorateAddress> address_space_cast(ElementType*);

// Deduction guides
template <typename T, int Dimensions, access::placeholder IsPlaceholder>
multi_ptr(accessor<T, Dimensions, access_mode::read, target::device, IsPlaceholder>)
    -> multi_ptr<const T, access::address_space::global_space, access::decorated::no>;

template <typename T, int Dimensions, access::placeholder IsPlaceholder>
multi_ptr(accessor<T, Dimensions, access_mode::write, target::device, IsPlaceholder>)
    -> multi_ptr<T, access::address_space::global_space, access::decorated::no>;

template <typename T, int Dimensions, access::placeholder IsPlaceholder>
multi_ptr(accessor<T, Dimensions, access_mode::read_write, target::device, IsPlaceholder>)
    -> multi_ptr<T, access::address_space::global_space, access::decorated::no>;

template <typename T, int Dimensions, access::placeholder IsPlaceholder>
multi_ptr(accessor<T, Dimensions, access_mode::read, target::constant_buffer, IsPlaceholder>)
    -> multi_ptr<const T, access::address_space::constant_space, access::decorated::no>;

template <typename T, int Dimensions, access_mode Mode, access::placeholder IsPlaceholder>
multi_ptr(accessor<T, Dimensions, Mode, target::local, IsPlaceholder>)
    -> multi_ptr<T, access::address_space::local_space, access::decorated::no>;

template <typename T, int Dimensions>
multi_ptr(local_accessor<T, Dimensions>)
    -> multi_ptr<T, access::address_space::local_space, access::decorated::no>;

} // namespace sycl
Table 61. Constructors of the SYCL multi_ptr class template
Constructor Description
multi_ptr()

Default constructor.

multi_ptr(const multi_ptr&)

Copy constructor.

multi_ptr(multi_ptr&&)

Move constructor.

explicit
multi_ptr(multi_ptr<ElementType, Space,
                    access::decorated::yes>::pointer)

Constructor that takes as an argument a decorated pointer.

multi_ptr(std::nullptr_t)

Constructor from a nullptr.

template <typename AccDataT, int Dimensions,
          access_mode Mode,
          access::placeholder IsPlaceholder>
multi_ptr(accessor<AccDataT, Dimensions, Mode,
                   target::device, IsPlaceholder>)

Available only when: (Space == access::address_space::global_space || Space == access::address_space::generic_space) && (std::is_void_v<ElementType> || std::is_same_v<std::remove_const_t<ElementType>, std::remove_const_t<AccDataT>>) && (std::is_const_v<ElementType> || !std::is_const_v<accessor<AccDataT, Dimensions, Mode, target::device, IsPlaceholder>::value_type>).

Constructs a multi_ptr from an accessor of target::device.

This constructor may only be called from within a command.

template <typename AccDataT, int Dimensions>
multi_ptr(local_accessor<AccDataT, Dimensions>)

Available only when: (Space == access::address_space::local_space || Space == access::address_space::generic_space) && (std::is_void_v<ElementType> || std::is_same_v<std::remove_const_t<ElementType>, std::remove_const_t<AccDataT>>) && (std::is_const_v<ElementType> || !std::is_const_v<AccDataT>).

Constructs a multi_ptr from a local_accessor.

This constructor may only be called from within a command.

template <typename AccDataT, int Dimensions,
          access_mode Mode,
          access::placeholder IsPlaceholder>
multi_ptr(accessor<AccDataT, Dimensions, Mode,
                   target::local, IsPlaceholder>)

Deprecated in SYCL 2020. Use the overload with local_accessor instead.

Available only when: (Space == access::address_space::local_space || Space == access::address_space::generic_space) && (std::is_void_v<ElementType> || std::is_same_v<std::remove_const_t<ElementType>, std::remove_const_t<AccDataT>>) && (std::is_const_v<ElementType> || !std::is_const_v<AccDataT>).

Constructs a multi_ptr from an accessor of target::local.

This constructor may only be called from within a command.

template <typename ElementType,
          access::address_space Space,
          access::decorated DecorateAddress>
multi_ptr<ElementType, Space, DecorateAddress>
make_ptr(ElementType* pointer)

Deprecated in SYCL 2020. Use address_space_cast instead.

Global function to create a multi_ptr instance depending on the address space of the pointer argument. An implementation must return nullptr if the run-time value of pointer is not compatible with Space, and must issue a compile-time diagnostic if the deduced address space is not compatible with Space.

template <access::address_space Space,
          access::decorated DecorateAddress,
          typename ElementType>
multi_ptr<ElementType, Space, DecorateAddress>
address_space_cast(ElementType* pointer)

Global function to create a multi_ptr instance from pointer, using the address space and decoration specified via the Space and DecorateAddress template arguments.

An implementation must return nullptr if the run-time value of pointer is not compatible with Space, and must issue a compile-time diagnostic if the deduced address space for pointer is not compatible with Space.

Table 62. Operators of multi_ptr class
Operators Description
multi_ptr& operator=(const multi_ptr&)

Copy assignment operator.

multi_ptr& operator=(multi_ptr&&)

Move assignment operator.

multi_ptr& operator=(std::nullptr_t)

Assigns nullptr to the multi_ptr.

template <access::address_space AS,
          access::decorated IsDecorated>
multi_ptr&
operator=(const multi_ptr<value_type, AS, IsDecorated>&)

Available only when: (Space == access::address_space::generic_space && AS != access::address_space::constant_space).

Assigns the value of the right hand side multi_ptr into the generic_ptr.

template<access::address_space AS,
         access::decorated IsDecorated>
multi_ptr&
operator=(multi_ptr<value_type, AS, IsDecorated>&&)

Available only when: (Space == access::address_space::generic_space && AS != access::address_space::constant_space).

Move the value of the right hand side multi_ptr into the generic_ptr.

reference operator[](std::ptrdiff_t i) const

Available only when: (!std::is_void_v<value_type>).

Returns a reference to the i-th pointed value. The value i can be negative.

pointer operator->() const

Available only when: (!std::is_void_v<value_type>).

Returns the underlying pointer.

reference operator*() const

Available only when: (!std::is_void_v<value_type>).

Returns a reference to the pointed value.

operator pointer() const

Implicit conversion to the underlying pointer type. Deprecated: The member function get should be used instead

template <access::decorated IsDecorated>
explicit
operator multi_ptr<value_type,
                   access::address_space::private_space,
                   IsDecorated>() const

Available only when: (Space == access::address_space::generic_space).

Conversion from generic_ptr to private_ptr. The result is undefined if the pointer does not address the private address space.

template <access::decorated IsDecorated>
explicit
operator multi_ptr<const value_type,
                   access::address_space::private_space,
                   IsDecorated>() const

Available only when: (Space == access::address_space::generic_space).

Conversion from generic_ptr to private_ptr of const data. The result is undefined if the pointer does not address the private address space.

template <access::decorated IsDecorated>
explicit
operator multi_ptr<value_type,
                   access::address_space::global_space,
                   IsDecorated>() const

Available only when: (Space == access::address_space::generic_space).

Conversion from generic_ptr to global_ptr. The result is undefined if the pointer does not address the global address space.

template <access::decorated IsDecorated>
explicit
operator multi_ptr<const value_type,
                   access::address_space::global_space,
                   IsDecorated>() const

Available only when: (Space == access::address_space::generic_space).

Conversion from generic_ptr to global_ptr of const data. The result is undefined if the pointer does not address the global address space.

template <access::decorated IsDecorated>
explicit
operator multi_ptr<value_type,
                   access::address_space::local_space,
                   IsDecorated>() const

Available only when: (Space == access::address_space::generic_space).

Conversion from generic_ptr to local_ptr. The result is undefined if the pointer does not address the local address space.

template <access::decorated IsDecorated>
explicit
operator multi_ptr<const value_type,
                   access::address_space::local_space,
                   IsDecorated>() const

Available only when: (Space == access::address_space::generic_space).

Conversion from generic_ptr to local_ptr of const data. The result is undefined if the pointer does not address the local address space.

template <access::decorated IsDecorated>
operator multi_ptr<void, Space, IsDecorated>() const

Available only when: (!std::is_void_v<value_type> && !std::is_const_v<value_type>).

Implicit conversion to a multi_ptr of type void.

template <access::decorated IsDecorated>
operator multi_ptr<const void, Space, IsDecorated>() const

Available only when: (!std::is_void_v<value_type> && std::is_const_v<value_type>).

Implicit conversion to a multi_ptr of type const void.

template <access::decorated IsDecorated>
operator multi_ptr<const value_type, Space,
                   IsDecorated>() const

Implicit conversion to a multi_ptr of type const value_type.

operator multi_ptr<value_type, Space,
                   access::decorated::no>() const

Available only when: (is_decorated == true).

Implicit conversion to the equivalent multi_ptr object that does not expose decorated pointers or references.

operator multi_ptr<value_type, Space,
                   access::decorated::yes>() const

Available only when: (is_decorated == false).

Implicit conversion to the equivalent multi_ptr object that exposes decorated pointers and references.

Table 63. Member functions of multi_ptr class
Member function Description
pointer get() const

Returns the underlying pointer. Whether the pointer is decorated depends on the value of DecorateAddress.

__unspecified__* get_decorated() const

Returns the underlying pointer decorated by the address space that it addresses. Note that the support involves implementation-defined device compiler extensions.

std::add_pointer_t<value_type> get_raw() const

Returns the underlying pointer, always undecorated.

void prefetch(std::size_t numElements) const

Available only when: Space == access::address_space::global_space.

Prefetches a number of elements specified by numElements into the global memory cache. This operation is an implementation-defined optimization and does not effect the functional behavior of the SYCL kernel function.

Table 64. Hidden friend functions of the multi_ptr class
Hidden friend function Description
reference operator*(const multi_ptr& mp)

Available only when: (!std::is_void_v<ElementType>).

Operator that returns a reference to the value_type of mp.

multi_ptr& operator++(multi_ptr& mp)

Available only when: (!std::is_void_v<ElementType>).

Increments mp by 1 and returns mp.

multi_ptr operator++(multi_ptr& mp, int)

Available only when: (!std::is_void_v<ElementType>).

Increments mp by 1 and returns a new multi_ptr with the value of the original mp.

multi_ptr& operator--(multi_ptr& mp)

Available only when: (!std::is_void_v<ElementType>).

Decrements mp by 1 and returns mp.

multi_ptr operator--(multi_ptr& mp, int)

Available only when: (!std::is_void_v<ElementType>).

Decrements mp by 1 and returns a new multi_ptr with the value of the original mp.

multi_ptr& operator+=(multi_ptr& lhs, difference_type r)

Available only when: (!std::is_void_v<ElementType>).

Moves mp forward by r and returns lhs.

multi_ptr& operator-=(multi_ptr& lhs, difference_type r)

Available only when: (!std::is_void_v<ElementType>).

Moves mp backward by r and returns lhs.

multi_ptr operator+(const multi_ptr& lhs, difference_type r)

Available only when: (!std::is_void_v<ElementType>).

Creates a new multi_ptr that points r forward compared to lhs.

multi_ptr operator-(const multi_ptr& lhs, difference_type r)

Available only when: (!std::is_void_v<ElementType>).

Creates a new multi_ptr that points r backward compared to lhs.

bool operator==(const multi_ptr& lhs, const multi_ptr& rhs)

Comparison operator == for multi_ptr class.

bool operator!=(const multi_ptr& lhs, const multi_ptr& rhs)

Comparison operator != for multi_ptr class.

bool operator<(const multi_ptr& lhs, const multi_ptr& rhs)

Comparison operator < for multi_ptr class.

bool operator>(const multi_ptr& lhs, const multi_ptr& rhs)

Comparison operator > for multi_ptr class.

bool operator<=(const multi_ptr& lhs, const multi_ptr& rhs)

Comparison operator <= for multi_ptr class.

bool operator>=(const multi_ptr& lhs, const multi_ptr& rhs)

Comparison operator >= for multi_ptr class.

bool operator==(const multi_ptr& lhs, std::nullptr_t)

Comparison operator == for multi_ptr class with a std::nullptr_t.

bool operator!=(const multi_ptr& lhs, std::nullptr_t)

Comparison operator != for multi_ptr class with a std::nullptr_t.

bool operator<(const multi_ptr& lhs, std::nullptr_t)

Comparison operator < for multi_ptr class with a std::nullptr_t.

bool operator>(const multi_ptr& lhs, std::nullptr_t)

Comparison operator > for multi_ptr class with a std::nullptr_t.

bool operator<=(const multi_ptr& lhs, std::nullptr_t)

Comparison operator <= for multi_ptr class with a std::nullptr_t.

bool operator>=(const multi_ptr& lhs, std::nullptr_t)

Comparison operator >= for multi_ptr class with a std::nullptr_t.

bool operator==(std::nullptr_t, const multi_ptr& rhs)

Comparison operator == for multi_ptr class with a std::nullptr_t.

bool operator!=(std::nullptr_t, const multi_ptr& rhs)

Comparison operator != for multi_ptr class with a std::nullptr_t.

bool operator<(std::nullptr_t, const multi_ptr& rhs)

Comparison operator < for multi_ptr class with a std::nullptr_t.

bool operator>(std::nullptr_t, const multi_ptr& rhs)

Comparison operator > for multi_ptr class with a std::nullptr_t.

bool operator<=(std::nullptr_t, const multi_ptr& rhs)

Comparison operator <= for multi_ptr class with a std::nullptr_t.

bool operator>=(std::nullptr_t, const multi_ptr& rhs)

Comparison operator >= for multi_ptr class with a std::nullptr_t.

The following is the overview of the legacy interface from 1.2.1 provided for the multi_ptr class.

  1
  2
  3
  4
  5
  6
  7
  8
  9
 10
 11
 12
 13
 14
 15
 16
 17
 18
 19
 20
 21
 22
 23
 24
 25
 26
 27
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
namespace sycl {

// Legacy interface, inherited from 1.2.1.
template <typename ElementType, access::address_space Space>
class [[deprecated]] multi_ptr<ElementType, Space, access::decorated::legacy> {
 public:
  using value_type = ElementType;
  using element_type = ElementType;
  using difference_type = std::ptrdiff_t;

  // Implementation defined pointer and reference types that correspond to
  // SYCL/OpenCL interoperability types for OpenCL C functions.
  using pointer_t =
      multi_ptr<ElementType, Space, access::decorated::yes>::pointer;
  using const_pointer_t =
      multi_ptr<const ElementType, Space, access::decorated::yes>::pointer;
  using reference_t =
      multi_ptr<ElementType, Space, access::decorated::yes>::reference;
  using const_reference_t =
      multi_ptr<const ElementType, Space, access::decorated::yes>::reference;

  static constexpr access::address_space address_space = Space;

  // Constructors
  multi_ptr();
  multi_ptr(const multi_ptr&);
  multi_ptr(multi_ptr&&);
  multi_ptr(pointer_t);
  multi_ptr(ElementType*);
  multi_ptr(std::nullptr_t);
  ~multi_ptr();

  // Assignment and access operators
  multi_ptr& operator=(const multi_ptr&);
  multi_ptr& operator=(multi_ptr&&);
  multi_ptr& operator=(pointer_t);
  multi_ptr& operator=(ElementType*);
  multi_ptr& operator=(std::nullptr_t);
  friend ElementType& operator*(const multi_ptr& mp) { /* ... */
  }
  ElementType* operator->() const;

  // Available only when:
  //   (Space == access::address_space::global_space ||
  //    Space == access::address_space::generic_space) &&
  //   (std::is_same_v<std::remove_const_t<ElementType>, std::remove_const_t<AccDataT>>) &&
  //   (std::is_const_v<ElementType> ||
  //    !std::is_const_v<accessor<AccDataT, Dimensions, Mode, target::device,
  //                              IsPlaceholder>::value_type>)
  template <int Dimensions, access_mode Mode, access::placeholder IsPlaceholder>
  multi_ptr(
      accessor<ElementType, Dimensions, Mode, target::device, IsPlaceholder>);

  // Available only when:
  //   (Space == access::address_space::local_space ||
  //    Space == access::address_space::generic_space) &&
  //   (std::is_same_v<std::remove_const_t<ElementType>, std::remove_const_t<AccDataT>>) &&
  //   (std::is_const_v<ElementType> || !std::is_const_v<AccDataT>)
  template <int Dimensions, access_mode Mode, access::placeholder IsPlaceholder>
  multi_ptr(
      accessor<ElementType, Dimensions, Mode, target::local, IsPlaceholder>);

  // Available only when:
  //   (Space == access::address_space::local_space ||
  //    Space == access::address_space::generic_space) &&
  //   (std::is_same_v<std::remove_const_t<ElementType>, std::remove_const_t<AccDataT>>) &&
  //   (std::is_const_v<ElementType> || !std::is_const_v<AccDataT>)
  template <typename AccDataT, int Dimensions>
  multi_ptr(local_accessor<AccDataT, Dimensions>);

  // Only if Space == constant_space
  template <int Dimensions, access_mode Mode, access::placeholder IsPlaceholder>
  multi_ptr(accessor<ElementType, Dimensions, Mode, target::constant_buffer,
                     IsPlaceholder>);

  // Returns the underlying OpenCL C pointer
  pointer_t get() const;

  std::add_pointer_t<value_type> get_raw() const;

  pointer_t get_decorated() const;

  // Implicit conversion to the underlying pointer type
  operator ElementType*() const;

  // Implicit conversion to a multi_ptr<void>
  // Available only when ElementType is not const-qualified
  operator multi_ptr<void, Space, access::decorated::legacy>() const;

  // Implicit conversion to a multi_ptr<const void>
  // Available only when ElementType is const-qualified
  operator multi_ptr<const void, Space, access::decorated::legacy>() const;

  // Implicit conversion to multi_ptr<const ElementType, Space>
  operator multi_ptr<const ElementType, Space, access::decorated::legacy>()
      const;

  // Arithmetic operators
  friend multi_ptr& operator++(multi_ptr& mp) { /* ... */
  }
  friend multi_ptr operator++(multi_ptr& mp, int) { /* ... */
  }
  friend multi_ptr& operator--(multi_ptr& mp) { /* ... */
  }
  friend multi_ptr operator--(multi_ptr& mp, int) { /* ... */
  }
  friend multi_ptr& operator+=(multi_ptr& lhs, difference_type r) { /* ... */
  }
  friend multi_ptr& operator-=(multi_ptr& lhs, difference_type r) { /* ... */
  }
  friend multi_ptr operator+(const multi_ptr& lhs,
                             difference_type r) { /* ... */
  }
  friend multi_ptr operator-(const multi_ptr& lhs,
                             difference_type r) { /* ... */
  }

  void prefetch(std::size_t numElements) const;

  friend bool operator==(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator!=(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator<(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator>(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator<=(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator>=(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }

  friend bool operator==(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator!=(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator<(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator>(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator<=(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator>=(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }

  friend bool operator==(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator!=(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator<(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator>(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator<=(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator>=(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
};

// Legacy interface, inherited from 1.2.1.
// Specialization of multi_ptr for void and const void
// VoidType can be either void or const void
template <access::address_space Space>
class [[deprecated]] multi_ptr<VoidType, Space, access::decorated::legacy> {
 public:
  using value_type = VoidType;
  using element_type = VoidType;
  using difference_type = std::ptrdiff_t;

  // Implementation defined pointer types that correspond to
  // SYCL/OpenCL interoperability types for OpenCL C functions
  using pointer_t = multi_ptr<VoidType, Space, access::decorated::yes>::pointer;
  using const_pointer_t =
      multi_ptr<const VoidType, Space, access::decorated::yes>::pointer;

  static constexpr access::address_space address_space = Space;

  // Constructors
  multi_ptr();
  multi_ptr(const multi_ptr&);
  multi_ptr(multi_ptr&&);
  multi_ptr(pointer_t);
  multi_ptr(VoidType*);
  multi_ptr(std::nullptr_t);
  ~multi_ptr();

  // Assignment operators
  multi_ptr& operator=(const multi_ptr&);
  multi_ptr& operator=(multi_ptr&&);
  multi_ptr& operator=(pointer_t);
  multi_ptr& operator=(VoidType*);
  multi_ptr& operator=(std::nullptr_t);

  // Available only when:
  //   (Space == access::address_space::global_space ||
  //    Space == access::address_space::generic_space) &&
  //   (std::is_const_v<VoidType> ||
  //    !std::is_const_v<accessor<ElementType, Dimensions, Mode, target::device,
  //                              IsPlaceholder>::value_type>)
  template <typename ElementType, int Dimensions, access_mode Mode>
  multi_ptr(accessor<ElementType, Dimensions, Mode, target::device>);

  // Available only when:
  //   (Space == access::address_space::local_space ||
  //    Space == access::address_space::generic_space) &&
  //   (std::is_const_v<VoidType> || !std::is_const_v<ElementType>)
  template <typename ElementType, int Dimensions, access_mode Mode>
  multi_ptr(accessor<ElementType, Dimensions, Mode, target::local>);

  // Available only when:
  //   (Space == access::address_space::local_space ||
  //    Space == access::address_space::generic_space) &&
  //   (std::is_const_v<VoidType> || !std::is_const_v<ElementType>)
  template <typename AccDataT, int Dimensions>
  multi_ptr(local_accessor<AccDataT, Dimensions>);

  // Only if Space == access::address_space::constant_space
  template <typename ElementType, int Dimensions, access_mode Mode>
  multi_ptr(accessor<ElementType, Dimensions, Mode, target::constant_buffer>);

  // Returns the underlying OpenCL C pointer
  pointer_t get() const;

  std::add_pointer_t<value_type> get_raw() const;

  pointer_t get_decorated() const;

  // Implicit conversion to the underlying pointer type
  operator VoidType*() const;

  // Explicit conversion to a multi_ptr<ElementType>
  // If VoidType is const, ElementType must be as well
  template <typename ElementType>
  explicit
  operator multi_ptr<ElementType, Space, access::decorated::legacy>() const;

  // Implicit conversion to multi_ptr<const void, Space>
  operator multi_ptr<const void, Space, access::decorated::legacy>() const;

  friend bool operator==(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator!=(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator<(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator>(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator<=(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator>=(const multi_ptr& lhs, const multi_ptr& rhs) { /* ... */
  }

  friend bool operator==(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator!=(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator<(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator>(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator<=(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }
  friend bool operator>=(const multi_ptr& lhs, std::nullptr_t) { /* ... */
  }

  friend bool operator==(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator!=(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator<(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator>(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator<=(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
  friend bool operator>=(std::nullptr_t, const multi_ptr& rhs) { /* ... */
  }
};

} // namespace sycl
4.7.7.2. Explicit pointer aliases

SYCL provides aliases to the multi_ptr class template (see Section 4.7.7.1) for each specialization of access::address_space.

A synopsis of the SYCL multi_ptr class template aliases is provided below.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
namespace sycl {

template <typename ElementType, access::address_space Space,
          access::decorated IsDecorated>
class multi_ptr;

// Template specialization aliases for different pointer address spaces

template <typename ElementType,
          access::decorated IsDecorated = access::decorated::legacy>
using global_ptr =
    multi_ptr<ElementType, access::address_space::global_space, IsDecorated>;

template <typename ElementType,
          access::decorated IsDecorated = access::decorated::legacy>
using local_ptr =
    multi_ptr<ElementType, access::address_space::local_space, IsDecorated>;

// Deprecated in SYCL 2020
template <typename ElementType>
using constant_ptr =
    multi_ptr<ElementType, access::address_space::constant_space,
              access::decorated::legacy>;

template <typename ElementType,
          access::decorated IsDecorated = access::decorated::legacy>
using private_ptr =
    multi_ptr<ElementType, access::address_space::private_space, IsDecorated>;

template <typename ElementType,
          access::decorated IsDecorated = access::decorated::legacy>
using generic_ptr =
    multi_ptr<ElementType, access::address_space::generic_space, IsDecorated>;

// Template specialization aliases for different pointer address spaces.
// The interface exposes non-decorated pointer while keeping the
// address space information internally.

template <typename ElementType>
using raw_global_ptr =
    multi_ptr<ElementType, access::address_space::global_space,
              access::decorated::no>;

template <typename ElementType>
using raw_local_ptr = multi_ptr<ElementType, access::address_space::local_space,
                                access::decorated::no>;

template <typename ElementType>
using raw_private_ptr =
    multi_ptr<ElementType, access::address_space::private_space,
              access::decorated::no>;

template <typename ElementType>
using raw_generic_ptr =
    multi_ptr<ElementType, access::address_space::generic_space,
              access::decorated::no>;

// Template specialization aliases for different pointer address spaces.
// The interface exposes decorated pointer.

template <typename ElementType>
using decorated_global_ptr =
    multi_ptr<ElementType, access::address_space::global_space,
              access::decorated::yes>;

template <typename ElementType>
using decorated_local_ptr =
    multi_ptr<ElementType, access::address_space::local_space,
              access::decorated::yes>;

template <typename ElementType>
using decorated_private_ptr =
    multi_ptr<ElementType, access::address_space::private_space,
              access::decorated::yes>;

template <typename ElementType>
using decorated_generic_ptr =
    multi_ptr<ElementType, access::address_space::generic_space,
              access::decorated::yes>;

} // namespace sycl

4.7.8. Image samplers

The SYCL image_sampler struct contains a configuration for sampling a sampled_image. The members of this struct are defined by the following tables.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
namespace sycl {

enum class addressing_mode : /* unspecified */ {
  mirrored_repeat,
  repeat,
  clamp_to_edge,
  clamp,
  none
};

enum class filtering_mode : /* unspecified */ { nearest, linear };

enum class coordinate_normalization_mode : /* unspecified */ {
  normalized,
  unnormalized
};

struct image_sampler {
  addressing_mode addressing;
  coordinate_normalization_mode coordinate;
  filtering_mode filtering;
};

} // namespace sycl
Table 65. Addressing modes description
addressing_mode Description
mirrored_repeat

Out of range coordinates will be flipped at every integer junction. This addressing mode can only be used with normalized coordinates. If normalized coordinates are not used, this addressing mode may generate image coordinates that are undefined.

repeat

Out of range image coordinates are wrapped to the valid range. This addressing mode can only be used with normalized coordinates. If normalized coordinates are not used, this addressing mode may generate image coordinates that are undefined.

clamp_to_edge

Out of range image coordinates are clamped to the extent.

clamp

Out of range image coordinates will return a border color.

none

For this addressing mode the programmer guarantees that the image coordinates used to sample elements of the image refer to a location inside the image; otherwise the results are undefined.

Table 66. Filtering modes description
filtering_mode Description
nearest

Chooses a color of nearest pixel.

linear

Performs a linear sampling of adjacent pixels.

Table 67. Coordinate normalization modes description
coordinate_normalization_mode Description
normalized

Normalizes image coordinates.

unnormalized

Does not normalize image coordinates.

4.8. Unified shared memory (USM)

This section describes properties and routines for pointer-based memory management interfaces in SYCL. These routines augment, rather than replace, the buffer-based interfaces in SYCL.

Unified Shared Memory (USM) provides a pointer-based alternative to the buffer programming model. USM enables:

  • Easier integration into existing code bases by representing allocations as pointers rather than buffers, with full support for pointer arithmetic into allocations.

  • Fine-grain control over ownership and accessibility of allocations, to optimally choose between performance and programmer convenience.

  • A simpler programming model, by automatically migrating some allocations between SYCL devices and the host.

To show the differences with the example from Section 3.2, the following source code example shows how shared memory can be used between host and device:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
#include <iostream>
#include <sycl/sycl.hpp>
using namespace sycl;  // (optional) avoids need for "sycl::" before SYCL names

int main() {
  //  Create a default queue to enqueue work to the default device
  queue myQueue;

  // Allocate shared memory bound to the device and context associated to the
  // queue Replacing malloc_shared with malloc_host would yield a correct
  // program that allocated device-visible memory on the host.
  int* data = sycl::malloc_shared<int>(1024, myQueue);

  myQueue.parallel_for(1024, [=](id<1> idx) {
    // Initialize each buffer element with its own rank number starting at 0
    data[idx] = idx;
  });  // End of the kernel function

  // Explicitly wait for kernel execution since there is no accessor involved
  myQueue.wait();

  // Print result
  for (int i = 0; i < 1024; i++)
    std::cout << "data[" << i << "] = " << data[i] << std::endl;

  return 0;
}

By comparison, the following source code example uses less capable device memory, which requires an explicit copy between the device and the host:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
#include <iostream>
#include <sycl/sycl.hpp>
using namespace sycl;  // (optional) avoids need for "sycl::" before SYCL names

int main() {
  // Create a default queue to enqueue work to the default device
  queue myQueue;

  // Allocate device USM, using the device and context associated with the queue
  int* data = sycl::malloc_device<int>(1024, myQueue);

  myQueue.parallel_for(1024, [=](id<1> idx) {
    // Initialize each buffer element with its own rank number starting at 0
    data[idx] = idx;
  });  // End of the kernel function

  // Explicitly wait for kernel execution since there is no accessor involved
  myQueue.wait();

  // Create an array to receive the device content
  int hostData[1024];
  // Receive the content from the device
  myQueue.memcpy(hostData, data, 1024 * sizeof(int));
  // Wait for the copy to complete
  myQueue.wait();

  // Print result
  for (int i = 0; i < 1024; i++)
    std::cout << "hostData[" << i << "] = " << hostData[i] << std::endl;

  return 0;
}

4.8.1. Unified addressing

The following guarantees apply to the pointer values of USM allocations for objects with overlapping lifetimes:

  • A USM pointer is different from all non-USM memory addresses on the host.

  • A USM host memory pointer allocated from context C has a value that is different from all other USM host memory pointers (regardless of context), different from all USM shared memory pointers (regardless of context), and different from all USM device memory pointers allocated from context C.

  • A USM shared memory pointer allocated from context C has a value that is different from all other USM shared memory pointers (regardless of context), different from all USM host memory pointers (regardless of context), and different from all USM device memory pointers allocated from context C.

  • A USM device memory pointer allocated from context C has a value that is different from all other USM pointers allocated from context C.

4.8.2. Kinds of unified shared memory

USM is a capability that, when available, provides the ability to create allocations that are visible to both host and device(s). USM builds upon Unified Addressing to define a shared address space where pointer values in this space always refer to the same location in memory. USM defines three types of memory allocations described in Table 68.

Table 68. Type of USM allocations
USM allocation type Description

host

Allocations in host memory that are accessible by a device

device

Allocations in device memory that are not accessible by the host

shared

Allocations in shared memory that are accessible by both host and device

The following enum is used to refer to the different types of allocations inside of a SYCL program:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
namespace sycl {
namespace usm {

enum class alloc : /* unspecified */ {
  host,
  device,
  shared,
  unknown
};

}
}

USM is an optional feature which may not be supported by all devices, and devices that support USM may not support all types of USM allocation. A SYCL application can use the device::has() function to determine the level of USM support for a device. See Section 4.6.4.5 for more details.

The characteristics of USM allocations are summarized in Table 69.

Table 69. Characteristics of the different kinds of USM allocation
Allocation Type Initial Location Accessible By Migratable To

device

device

host

No

host

No

device

Yes

device

N/A

Another device

Optional (P2P)

Another device

No

host

host

host

Yes

host

N/A

Any device

Yes

device

No

shared

Unspecified

host

Yes

host

Yes

device

Yes

device

Yes

Another device

Optional

Another device

Optional

Each USM allocation has an associated SYCL context, and any access to that memory must use the same context. Specifically, any SYCL kernel function that dereferences a pointer to a USM allocation must be submitted to a queue that was constructed with the same context that was used to allocate that memory. The explicit memory operation commands that take USM pointers have a similar restriction. (See Section 4.9.4.3 for details.) Violations of these requirements result in undefined behavior.

There are no similar restrictions for dereferencing a USM pointer in a host task. This is legal regardless of which queue the host task was submitted to so long as the USM pointer is accessible on the host.

Each type of USM allocation has different rules for where that memory is accessible. Attempting to dereference a USM pointer on the host or on a device in violation of these rules results in undefined behavior. Passing a USM pointer to one of the explicit memory functions where the pointer is not accessible to the device generally results in undefined behavior. See Section 4.9.4.3 for the exact rules.

Device allocations are used for explicitly managing device memory. Programmers directly allocate device memory and explicitly copy data between host memory and a device allocation. Device allocations are obtained through SYCL device USM allocation routines instead of system allocation routines like std::malloc or C++ new. Device allocations are not accessible on the host, but the pointer values remain consistent on account of Unified Addressing. The size of device allocations will be limited by the amount of memory in a device. Support for device allocations on a specific device can be queried through aspect::usm_device_allocations.

Device allocations must be explicitly copied between the host and a device. The member functions to copy and initialize data are found in Section 4.6.5.3 and Table 102, and these functions may be used on device allocations if a device supports aspect::usm_device_allocations.

Host allocations allow devices to directly read and write host memory inside of a kernel. This can be useful for several reasons, such as when the overhead of moving a small amount of data is not worth paying over the cost of a remote access or when the size of a data set exceeds the size of a device’s memory. Host allocations must also be obtained using SYCL routines instead of system allocation routines. While a device may remotely read and write a host allocation, the allocation does not migrate to the device - it remains in host memory. Users should take care to properly synchronize access to host allocations between host execution and kernels. The total size of host allocations will be limited by the amount of pinnable-memory on the host on most systems. Support for host allocations on a specific device can be queried through aspect::usm_host_allocations. Support for atomic modification of host allocations on a specific device can be queried through aspect::usm_atomic_host_allocations.

Shared allocations implicitly share data between the host and devices. Data may move to where it is being used without the programmer explicitly informing the runtime. It is up to the runtime and backends to make sure that a shared allocation is available where it is used. Shared allocations must also be obtained using SYCL allocation routines instead of the system allocator. The maximum size of a shared allocation on a specific device, and the total size of all shared allocations in a context, are implementation-defined. Support for shared allocations on a specific device can be queried through aspect::usm_shared_allocations.

Not all devices may support concurrent access of a shared allocation with the host. If a device does not support this, host execution and device code must take turns accessing the allocation, so the host must not access a shared allocation while a kernel is executing. Host access to a shared allocation which is also accessed by an executing kernel on a device that does not support concurrent access results in undefined behavior. If a device does support concurrent access, both the host and and the device may atomically modify the same data inside an allocation. Allocations, or pieces of allocations, are now free to migrate to different devices in the same context that also support this capability. Additionally, many devices that support concurrent access may support a working set of shared allocations larger than device memory. Users may query whether a device supports concurrent access with atomic modification of shared allocations through the aspect aspect::usm_atomic_shared_allocations. See Section 4.6.4.5 for more details.

Performance hints for shared allocations may be specified by the user by enqueuing prefetch operations on a device. These operations inform the SYCL runtime that the specified shared allocation is likely to be accessed on the device in the future, and that it is free to migrate the allocation to the device. More about prefetch is found in Section 4.6.5.3 and Table 102. If a device supports concurrent access to shared allocations, then prefetch operations may be overlapped with kernel execution.

Additionally, users may use the mem_advise member function to annotate shared allocations with advice. Valid advice is defined by the device and its associated backend. See Section 4.6.5.3 and Table 102 for more information.

In the most capable systems, users do not need to use SYCL USM allocation functions to create shared allocations. The system allocator (malloc/new) may instead be used. Likewise, std::free and delete are used instead of sycl::free. Note that host and device allocations are unaffected by this change and must still be allocated using their respective USM functions in order to guarantee their behavior. Users may query the device to determine if system allocations are supported for use on the device, through aspect::usm_system_allocations.

4.8.3. USM allocations

USM provides several allocation functions. These functions accept a property_list parameter, which is provided for future extensibility. The core SYCL specification does not yet define any USM allocation properties.

Some of the allocation functions take an explicit alignment parameter. Like std::aligned_alloc, these functions return nullptr if the alignment is not supported by the implementation. Some of the allocation functions are templated on the allocated type T and some are not. The following table specifies the alignment guarantees for each category.

Table 70. Alignment guarantees of USM allocation functions
Category Alignment guarantee

No alignment parameter
Not templated on allocation type

Pointer is suitably aligned for any object with fundamental alignment whose size is less than or equal to the requested allocation size.

No alignment parameter
Templated on allocation type T

Pointer is suitably aligned for an object of type T.

Alignment parameter alignment specified
Not templated on allocation type

Pointer is suitably aligned for any object with fundamental alignment whose size is less than or equal to the requested allocation size or it is aligned to the specified alignment, whichever is greater.

Alignment parameter alignment specified
Templated on allocation type T

Pointer is suitably aligned for an object of type T or it is aligned to the specified alignment, whichever is greater.

4.8.3.1. C++ allocator interface

SYCL defines an allocator class named usm_allocator that satisfies the C++ named requirement Allocator. The AllocKind template parameter can be either usm::alloc::host or usm::alloc::shared, causing the allocator to make either host USM allocations or shared USM allocations.

There is no specialization for usm::alloc::device because an Allocator is required to allocate memory that is accessible on the host.

The usm_allocator class has a template argument Alignment, which specifies the minimum alignment for memory that it allocates. This alignment is used even if the allocator is rebound to a different type. Memory allocated by this allocator is suitably aligned for objects of its underlying value_type or at the alignment specified by Alignment, whichever is greater.

A synopsis of the usm_allocator class is provided below. The constructors are listed in Table 71.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
template <typename T, usm::alloc AllocKind, std::size_t Alignment = 0>
class usm_allocator {
public:
  using value_type = T;
  using propagate_on_container_copy_assignment = std::true_type;
  using propagate_on_container_move_assignment = std::true_type;
  using propagate_on_container_swap = std::true_type;

public:
  template <typename U> struct rebind {
    typedef usm_allocator<U, AllocKind, Alignment> other;
  };

  usm_allocator() = delete;
  usm_allocator(const context& syclContext,
                const device& syclDevice,
                const property_list& propList = {});
  usm_allocator(const queue& syclQueue,
                const property_list& propList = {});
  usm_allocator(const usm_allocator& other);
  usm_allocator(usm_allocator&&) noexcept;
  usm_allocator& operator=(const usm_allocator&);
  usm_allocator& operator=(usm_allocator&&);

  template <class U>
  usm_allocator(usm_allocator<U, AllocKind, Alignment> const&) noexcept;

  /// Allocate memory
  T* allocate(std::size_t count);

  /// Deallocate memory
  void deallocate(T* Ptr, std::size_t count);

  /// Equality Comparison
  ///
  /// Allocators only compare equal if they are of the same USM kind, alignment,
  /// context, and device
  template <class U, usm::alloc AllocKindU, std::size_t AlignmentU>
  friend bool operator==(const usm_allocator<T, AllocKind, Alignment>&,
                         const usm_allocator<U, AllocKindU, AlignmentU>&);

  /// Inequality Comparison
  /// Allocators only compare unequal if they are not of the same USM kind, alignment,
  /// context, or device
  template <class U, usm::alloc AllocKindU, std::size_t AlignmentU>
  friend bool operator!=(const usm_allocator<T, AllocKind, Alignment>&,
                         const usm_allocator<U, AllocKindU, AlignmentU>&);
};
Table 71. Constructors of the usm_allocator class
Constructor Description
usm_allocator(const context& syclContext, const device& syclDevice,
              const property_list& propList = {})

Constructs a usm_allocator instance that allocates USM for the provided context and device.

If AllocKind is usm::alloc::host, this constructor throws a synchronous exception with the errc::feature_not_supported error code if no device in syclContext has aspect::usm_host_allocations. The syclDevice is ignored for this allocation kind.

If AllocKind is usm::alloc::shared, this constructor throws a synchronous exception with the errc::feature_not_supported error code if the syclDevice does not have aspect::usm_shared_allocations. The syclDevice must either be contained by syclContext or it must be a descendent device of some device that is contained by that context, otherwise this constructor throws a synchronous exception with the errc::invalid error code.

usm_allocator(const queue& syclQueue, const property_list& propList = {})

Simplified constructor form where syclQueue provides the device and context.

4.8.3.2. Device allocation functions

The functions in Table 72 allocate device USM. On success, these functions return a pointer to the newly allocated memory, which must eventually be deallocated with sycl::free in order to avoid a memory leak. If there are not enough resources to allocate the requested memory, these functions return nullptr.

When the allocation size is zero bytes (numBytes or count is zero), these functions behave in a manner consistent with C++ std::malloc. The value returned is unspecified in this case, and the returned pointer may not be used to access storage. If this pointer is not null, it must be passed to sycl::free to avoid a memory leak.

Table 72. Device USM Allocation Functions
Function Description
void* sycl::malloc_device(std::size_t numBytes, const device& syclDevice,
                          const context& syclContext,
                          const property_list& propList = {})

Returns a pointer to the newly allocated memory, which is allocated on syclDevice. The allocation size is specified in bytes. Throws a synchronous exception with the errc::feature_not_supported error code if the syclDevice does not have aspect::usm_device_allocations. The syclDevice must either be contained by syclContext or it must be a descendent device of some device that is contained by that context, otherwise this function throws a synchronous exception with the errc::invalid error code.

template <typename T>
T* sycl::malloc_device(std::size_t count, const device& syclDevice,
                       const context& syclContext,
                       const property_list& propList = {})

Returns a pointer to the newly allocated memory, which is allocated on syclDevice. The allocation size is specified in number of elements of type T. Throws a synchronous exception with the errc::feature_not_supported error code if the syclDevice does not have aspect::usm_device_allocations. The syclDevice must either be contained by syclContext or it must be a descendent device of some device that is contained by that context, otherwise this function throws a synchronous exception with the errc::invalid error code.

void* sycl::malloc_device(std::size_t numBytes, const queue& syclQueue,
                          const property_list& propList = {})

Simplified form where syclQueue provides the device and context.

template <typename T>
T* sycl::malloc_device(std::size_t count, const queue& syclQueue,
                       const property_list& propList = {})

Simplified form where syclQueue provides the device and context.

void* sycl::aligned_alloc_device(std::size_t alignment, std::size_t numBytes,
                                 const device& syclDevice,
                                 const context& syclContext,
                                 const property_list& propList = {})

Returns a pointer to the newly allocated memory, which is allocated on syclDevice. The allocation is specified in bytes and aligned according to alignment. Throws a synchronous exception with the errc::feature_not_supported error code if the syclDevice does not have aspect::usm_device_allocations. The syclDevice must either be contained by syclContext or it must be a descendent device of some device that is contained by that context, otherwise this function throws a synchronous exception with the errc::invalid error code.

template <typename T>
T* sycl::aligned_alloc_device(std::size_t alignment, std::size_t count,
                              const device& syclDevice,
                              const context& syclContext,
                              const property_list& propList = {})

Returns a pointer to the newly allocated memory, which is allocated on syclDevice. The allocation is specified in number of elements of type T and aligned according to alignment. Throws a synchronous exception with the errc::feature_not_supported error code if the syclDevice does not have aspect::usm_device_allocations. The syclDevice must either be contained by syclContext or it must be a descendent device of some device that is contained by that context, otherwise this function throws a synchronous exception with the errc::invalid error code.

void* sycl::aligned_alloc_device(std::size_t alignment, std::size_t numBytes,
                                 const queue& syclQueue,
                                 const property_list& propList = {})

Simplified form where syclQueue provides the device and context.

template <typename T>
T* sycl::aligned_alloc_device(std::size_t alignment, std::size_t count,
                              const queue& syclQueue,
                              const property_list& propList = {})

Simplified form where syclQueue provides the device and context.

4.8.3.3. Host allocation functions

The functions in Table 73 allocate host USM. On success, these functions return a pointer to the newly allocated memory, which must eventually be deallocated with sycl::free in order to avoid a memory leak. If there are not enough resources to allocate the requested memory, these functions return nullptr.

When the allocation size is zero bytes (numBytes or count is zero), these functions behave in a manner consistent with C++ std::malloc. The value returned is unspecified in this case, and the returned pointer may not be used to access storage. If this pointer is not null, it must be passed to sycl::free to avoid a memory leak.

Table 73. Host USM Allocation Functions
Function Description
void* sycl::malloc_host(std::size_t numBytes, const context& syclContext,
                        const property_list& propList = {})

Returns a pointer to the newly allocated memory. This allocation is specified in bytes. Throws a synchronous exception with the errc::feature_not_supported error code if no device in syclContext has aspect::usm_host_allocations.

template <typename T>
T* sycl::malloc_host(std::size_t count, const context& syclContext,
                     const property_list& propList = {})

Returns a pointer to the newly allocated memory. This allocation is specified in number of elements of type T. Throws a synchronous exception with the errc::feature_not_supported error code if no device in syclContext has aspect::usm_host_allocations.

void* sycl::malloc_host(std::size_t numBytes, const queue& syclQueue,
                        const property_list& propList = {})

Simplified form where syclQueue provides the context.

template <typename T>
T* sycl::malloc_host(std::size_t count, const queue& syclQueue,
                     const property_list& propList = {})

Simplified form where syclQueue provides the context.

void* sycl::aligned_alloc_host(std::size_t alignment, std::size_t numBytes,
                               const context& syclContext,
                               const property_list& propList = {})

Returns a pointer to the newly allocated memory. This allocation is specified in bytes and aligned according to alignment. Throws a synchronous exception with the errc::feature_not_supported error code if no device in syclContext has aspect::usm_host_allocations.

template <typename T>
T* sycl::aligned_alloc_host(std::size_t alignment, std::size_t count,
                            const context& syclContext,
                            const property_list& propList = {})

Returns a pointer to the newly allocated memory. This allocation is specified in elements of type T and aligned according to alignment. Throws a synchronous exception with the errc::feature_not_supported error code if no device in syclContext has aspect::usm_host_allocations.

void* sycl::aligned_alloc_host(std::size_t alignment, std::size_t numBytes,
                               const queue& syclQueue,
                               const property_list& propList = {})

Simplified form where syclQueue provides the context.

template <typename T>
T* sycl::aligned_alloc_host(std::size_t alignment, std::size_t count,
                               const queue& syclQueue,
                               const property_list& propList = {})

Simplified form where syclQueue provides the context.

4.8.3.4. Shared allocation functions

The functions in Table 74 allocate shared USM. On success, these functions return a pointer to the newly allocated memory, which must eventually be deallocated with sycl::free in order to avoid a memory leak. If there are not enough resources to allocate the requested memory, these functions return nullptr.

When the allocation size is zero bytes (numBytes or count is zero), these functions behave in a manner consistent with C++ std::malloc. The value returned is unspecified in this case, and the returned pointer may not be used to access storage. If this pointer is not null, it must be passed to sycl::free to avoid a memory leak.

Table 74. Shared USM Allocation Functions
Function Description
void* sycl::malloc_shared(std::size_t numBytes, const device& syclDevice,
                          const context& syclContext,
                          const property_list& propList = {})

Returns a pointer to the newly allocated memory, which is associated with syclDevice. This allocation is specified in bytes. Throws a synchronous exception with the errc::feature_not_supported error code if the syclDevice does not have aspect::usm_shared_allocations. The syclDevice must either be contained by syclContext or it must be a descendent device of some device that is contained by that context, otherwise this function throws a synchronous exception with the errc::invalid error code.

template <typename T>
T* sycl::malloc_shared(std::size_t count, const device& syclDevice,
                       const context& syclContext,
                       const property_list& propList = {})

Returns a pointer to the newly allocated memory, which is associated with syclDevice. This allocation is specified in number of elements of type T. Throws a synchronous exception with the errc::feature_not_supported error code if the syclDevice does not have aspect::usm_shared_allocations. The syclDevice must either be contained by syclContext or it must be a descendent device of some device that is contained by that context, otherwise this function throws a synchronous exception with the errc::invalid error code.

void* sycl::malloc_shared(std::size_t numBytes, const queue& syclQueue,
                          const property_list& propList = {})

Simplified form where syclQueue provides the device and context.

template <typename T>
T* sycl::malloc_shared(std::size_t count, const queue& syclQueue,
                       const property_list& propList = {})

Simplified form where syclQueue provides the device and context.

void* sycl::aligned_alloc_shared(std::size_t alignment, std::size_t numBytes,
                                 const device& syclDevice,
                                 const context& syclContext,
                                 const property_list& propList = {})

Returns a pointer to the newly allocated memory, which is associated with syclDevice. This allocation is specified in bytes and aligned according to alignment. Throws a synchronous exception with the errc::feature_not_supported error code if the syclDevice does not have aspect::usm_shared_allocations. The syclDevice must either be contained by syclContext or it must be a descendent device of some device that is contained by that context, otherwise this function throws a synchronous exception with the errc::invalid error code.

template <typename T>
T* sycl::aligned_alloc_shared(std::size_t alignment, std::size_t count,
                              const device& syclDevice,
                              const context& syclContext,
                              const property_list& propList = {})

Returns a pointer to the newly allocated memory, which is associated with syclDevice. This allocation is specified in number of elements of type T and aligned aligned according to alignment. Throws a synchronous exception with the errc::feature_not_supported error code if the syclDevice does not have aspect::usm_shared_allocations. The syclDevice must either be contained by syclContext or it must be a descendent device of some device that is contained by that context, otherwise this function throws a synchronous exception with the errc::invalid error code.

void* sycl::aligned_alloc_shared(std::size_t alignment, std::size_t numBytes,
                                 const queue& syclQueue,
                                 const property_list& propList = {})

Simplified form where syclQueue provides the device and context.

template <typename T>
T* sycl::aligned_alloc_shared(std::size_t alignment, std::size_t count,
                              const queue& syclQueue,
                              const property_list& propList = {})

Simplified form where syclQueue provides the device and context.

4.8.3.5. Parameterized allocation functions

The functions in Table 75 take a kind parameter that specifies the type of USM to allocate. When kind is usm::alloc::device, then the allocation device must have aspect::usm_device_allocations. When kind is usm::alloc::host, at least one device in the allocation context must have aspect::usm_host_allocations. When kind is usm::alloc::shared, the allocation device must have aspect::usm_shared_allocations. If these requirements are violated, the allocation function throws a synchronous exception with the errc::feature_not_supported error code.

On success, these functions return a pointer to the newly allocated memory, which must eventually be deallocated with sycl::free in order to avoid a memory leak. If there are not enough resources to allocate the requested memory, these functions return nullptr.

When the allocation size is zero bytes (numBytes or count is zero), these functions behave in a manner consistent with C++ std::malloc. The value returned is unspecified in this case, and the returned pointer may not be used to access storage. If this pointer is not null, it must be passed to sycl::free to avoid a memory leak.

Table 75. Parameterized USM Allocation Functions
Function Description
void* sycl::malloc(std::size_t numBytes, const device& syclDevice,
                   const context& syclContext, usm::alloc kind,
                   const property_list& propList = {})

Returns a pointer to the newly allocated memory of type kind. This allocation size is specified in bytes. The syclDevice parameter is ignored if kind is usm::alloc::host. If kind is not usm::alloc::host, syclDevice must either be contained by syclContext or it must be a descendent device of some device that is contained by that context, otherwise this function throws a synchronous exception with the errc::invalid error code.

template <typename T>
T* sycl::malloc(std::size_t count, const device& syclDevice,
                const context& syclContext, usm::alloc kind,
                const property_list& propList = {})

Returns a pointer to the newly allocated memory of type kind. This allocation size is specified in number of elements of type T. The syclDevice parameter is ignored if kind is usm::alloc::host. If kind is not usm::alloc::host, syclDevice must either be contained by syclContext or it must be a descendent device of some device that is contained by that context, otherwise this function throws a synchronous exception with the errc::invalid error code.

void* sycl::malloc(std::size_t numBytes, const queue& syclQueue, usm::alloc kind,
                   const property_list& propList = {})

Simplified form where syclQueue provides the context and any necessary device.

template <typename T>
T* sycl::malloc(std::size_t count, const queue& syclQueue, usm::alloc kind,
                const property_list& propList = {})

Simplified form where syclQueue provides the context and any necessary device.

void* sycl::aligned_alloc(std::size_t alignment, std::size_t numBytes,
                          const device& syclDevice, const context& syclContext,
                          usm::alloc kind, const property_list& propList = {})

Returns a pointer to the newly allocated memory of type kind. This allocation is specified in bytes and is aligned according to alignment. The syclDevice parameter is ignored if kind is usm::alloc::host. If kind is not usm::alloc::host, syclDevice must either be contained by syclContext or it must be a descendent device of some device that is contained by that context, otherwise this function throws a synchronous exception with the errc::invalid error code.

template <typename T>
T* sycl::aligned_alloc(std::size_t alignment, std::size_t count, const device& syclDevice,
                       const context& syclContext, usm::alloc kind,
                       const property_list& propList = {})

Returns a pointer to the newly allocated memory of type kind. This allocation is specified in number of elements of type T and is aligned according to alignment. The syclDevice parameter is ignored if kind is usm::alloc::host. If kind is not usm::alloc::host, syclDevice must either be contained by syclContext or it must be a descendent device of some device that is contained by that context, otherwise this function throws a synchronous exception with the errc::invalid error code.

void* sycl::aligned_alloc(std::size_t alignment, std::size_t numBytes,
                          const queue& syclQueue, usm::alloc kind,
                          const property_list& propList = {})

Simplified form where syclQueue provides the context and any necessary device.

template <typename T>
T* sycl::aligned_alloc(std::size_t alignment, std::size_t count, const queue& syclQueue,
                       usm::alloc kind, const property_list& propList = {})

Simplified form where syclQueue provides the context and any necessary device.

4.8.3.6. Memory deallocation functions
free
void free(void* ptr, const context& ctxt); (1)
void free(void* ptr, const queue& q);      (2)

Overload (1):

Preconditions:

  • ptr points to memory allocated against ctxt using one of the USM allocation routines, or is a null pointer;

  • ptr has not previously been deallocated; and

  • There are no in-progress or enqueued commands using the memory pointed to by ptr.

Effects: Causes the memory pointed to by ptr to be deallocated.

[Note: Whether free is blocking or non-blocking is unspecified. Applications should not rely on free for synchronization, nor assume that free cannot cause deadlocks.— end note]

Synchronization: A call to free that deallocates a region of memory synchronizes with any allocation call that allocates all or part of the same region of memory.

Remarks: If ptr is null, this function has no effect.

Overload (2):

Effects: Equivalent to return free(ptr, q.get_context());.

[Note: Although this overload accepts a queue argument, it does not submit a "free" command to the device; the queue argument is only used to determine the context associated with ptr.— end note]

4.8.4. Unified shared memory pointer queries

Since USM pointers look like raw C++ pointers, users cannot deduce what kind of USM allocation a given pointer may be from examining its type. However, two functions are defined that let users query the type of a USM allocation and, if applicable, the device on which it was allocated. These query functions are only supported on the host.

Table 76. USM Pointer Query Functions
Function Description
usm::alloc get_pointer_type(const void* ptr, const context& syclContext)

Returns the USM allocation type for ptr if ptr falls inside a valid USM allocation for the context syclContext. Returns usm::alloc::unknown if ptr does not point within a valid USM allocation from syclContext.

device get_pointer_device(const void* ptr, const context& syclContext)

Returns the device associated with the USM allocation. If ptr points within a device USM allocation or a shared USM allocation for the context syclContext, returns the same device that was passed when allocating the memory. If ptr points within a host USM allocation for the context syclContext, returns the first device in syclContext. Throws a synchronous exception with the errc::invalid error code if ptr does not point within a valid USM allocation from syclContext.

4.9. Expressing parallelism through kernels

4.9.1. Ranges and index space identifiers

The data parallelism of the SYCL kernel execution model requires instantiation of a parallel execution over a range of iteration space coordinates. To achieve this, SYCL exposes types to define the range of execution and to identify a given execution instance’s point in the iteration space.

The following types are defined: range, nd_range, id, item, h_item, nd_item and group.

When constructing multi-dimensional ids or ranges from integers, the elements are written such that the right-most element varies fastest in a linearization of the multi-dimensional space (see Section 3.11.1).

Table 77. Summary of types used to identify points in an index space, and ranges over which those points can vary
Type Description
id

A point within a range

range

Bounds over which an id may vary

item

Pairing of an id (specific point) and the range that it is bounded by

nd_range

Encapsulates both global and local (work-group size) ranges over which work-item ids will vary

nd_item

Encapsulates two items, one for global id and range, and one for local id and range

h_item

Index point queries within hierarchical parallelism (parallel_for_work_item). Encapsulates physical global and local ids and ranges, as well as a logical local id and range defined by hierarchical parallelism

group

Work-group queries within hierarchical parallelism (parallel_for_work_group), and exposes the parallel_for_work_item construct that identifies code to be executed by each work-item. Encapsulates work-group ids and ranges

4.9.1.1. range class

range<int Dimensions> is a 1D, 2D or 3D vector that defines the iteration domain of either a single work-group in a parallel dispatch, or the overall Dimensions of the dispatch. It can be constructed from integers.

The SYCL range class template provides the common by-value semantics (see Section 4.5.3).

A synopsis of the SYCL range class is provided below. The constructors, member functions and non-member functions of the SYCL range class are listed in Table 78, Table 79 and Table 80 respectively. The additional common special member functions and common member functions are listed in Section 4.5.3 in Table 9 and Table 10 respectively.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
namespace sycl {
template <int Dimensions = 1> class range {
 public:
  static constexpr int dimensions = Dimensions;

  range() noexcept;

  /* The following constructor is only available in the range class
   * specialization where: Dimensions==1 */
  range(std::size_t dim0) noexcept;
  /* The following constructor is only available in the range class
   * specialization where: Dimensions==2 */
  range(std::size_t dim0, std::size_t dim1) noexcept;
  /* The following constructor is only available in the range class
   * specialization where: Dimensions==3 */
  range(std::size_t dim0, std::size_t dim1, std::size_t dim2) noexcept;

  /* -- common interface members -- */

  std::size_t get(int dimension) const noexcept;
  std::size_t& operator[](int dimension) noexcept;
  std::size_t operator[](int dimension) const noexcept;

  std::size_t size() const noexcept;

  // OP is: +, -, *, /, %, <<, >>, &, |, ^, &&, ||, <, >, <=, >=
  friend range operatorOP(const range& lhs, const range& rhs) noexcept { /* ... */
  }

  // OP is: +, -, *, /, %, <<, >>, &, |, ^, &&, ||, <, >, <=, >=
  // Available only when std::is_integral_v<T> is true
  template <typename T>
  friend range operatorOP(const range& lhs, const T& rhs) noexcept { /* ... */
  }

  // OP is: +, -, *, /, %, <<, >>, &, |, ^, &&, ||, <, >, <=, >=
  // Available only when std::is_integral_v<T> is true
  template <typename T>
  friend range operatorOP(const T& lhs, const range& rhs) noexcept { /* ... */
  }

  // OP is: +=, -=, *=, /=, %=, <<=, >>=, &=, |=, ^=
  friend range& operatorOP(range& lhs, const range& rhs) noexcept { /* ... */
  }

  // OP is: +=, -=, *=, /=, %=, <<=, >>=, &=, |=, ^=
  // Available only when std::is_integral_v<T> is true
  template <typename T>
  friend range& operatorOP(range& lhs, const T& rhs) noexcept { /* ... */
  }

  // OP is unary +, -
  friend range operatorOP(const range& rhs) noexcept { /* ... */
  }

  // OP is prefix ++, --
  friend range& operatorOP(range& rhs) noexcept { /* ... */
  }

  // OP is postfix ++, --
  friend range operatorOP(range& lhs, int) noexcept { /* ... */
  }
};

// Deduction guides
range(std::size_t)->range<1>;
range(std::size_t, std::size_t)->range<2>;
range(std::size_t, std::size_t, std::size_t)->range<3>;

} // namespace sycl
Table 78. Constructors of the range class template
Constructor Description
range() noexcept;

Construct a SYCL range with the value 0 for each dimension.

range(std::size_t dim0) noexcept;

Construct a 1D range with value dim0. Only valid when the template parameter Dimensions is equal to 1.

range(std::size_t dim0, std::size_t dim1) noexcept;

Construct a 2D range with values dim0 and dim1. Only valid when the template parameter Dimensions is equal to 2.

range(std::size_t dim0, std::size_t dim1, std::size_t dim2) noexcept;

Construct a 3D range with values dim0, dim1 and dim2. Only valid when the template parameter Dimensions is equal to 3.

Table 79. Member functions of the range class template
Member function Description
std::size_t get(int dimension) const noexcept;

Return the value of the specified dimension of the range. Results in undefined behavior if dimension is not in the range [0, Dimensions).

std::size_t& operator[](int dimension) noexcept;

Return the l-value of the specified dimension of the range. Results in undefined behavior if dimension is not in the range [0, Dimensions).

std::size_t operator[](int dimension) const noexcept;

Return the value of the specified dimension of the range. Results in undefined behavior if dimension is not in the range [0, Dimensions).

std::size_t size() const noexcept;

Return the size of the range computed as dimension0*…​*dimensionN.

Table 80. Hidden friend functions of the SYCL range class template
Hidden friend function Description
range operatorOP(const range& lhs, const range& rhs) noexcept;

Where OP is: +, -, *, /, %, <<, >>, &, |, ^, &&, ||, <, >, <=, >=.

Constructs and returns a new instance of the SYCL range class template with the same dimensionality as lhs range, where each element of the new SYCL range instance is the result of an element-wise OP operator between each element of lhs range and each element of the rhs range. If the operator returns a bool, the result is then cast to std::size_t.

template <typename T>
range operatorOP(const range& lhs, const T& rhs) noexcept;

Where OP is: +, -, *, /, %, <<, >>, &, |, ^, &&, ||, <, >, <=, >=.

Constructs and returns a new instance of the SYCL range class template with the same dimensionality as lhs range, where each element of the new SYCL range instance is the result of an element-wise OP operator between each element of this SYCL range and the rhs integral type. If the operator returns a bool, the result is then cast to std::size_t.

template <typename T>
range operatorOP(const T& lhs, const range& rhs) noexcept;

Where OP is: +, -, *, /, %, <<, >>, &, |, ^, &&, ||, <, >, <=, >=.

Constructs and returns a new instance of the SYCL range class template with the same dimensionality as the rhs SYCL range, where each element of the new SYCL range instance is the result of an element-wise OP operator between the lhs integral type and each element of the rhs SYCL range. If the operator returns a bool, the result is then cast to std::size_t.

range& operatorOP(range& lhs, const range& rhs) noexcept;

Where OP is: +=, -=,*=, /=, %=, <<=, >>=, &=, |=, ^=.

Assigns each element of lhs range instance with the result of an element-wise OP operator between each element of lhs range and each element of the rhs range and returns lhs range. If the operator returns a bool, the result is then cast to std::size_t.

template <typename T>
range& operatorOP(range& lhs, const T& rhs) noexcept;

Where OP is: +=, -=,*=, /=, %=, <<=, >>=, &=, |=, ^=.

Assigns each element of lhs range instance with the result of an element-wise OP operator between each element of lhs range and the rhs integral type and returns lhs range. If the operator returns a bool, the result is then cast to std::size_t.

range operatorOP(const range& rhs) noexcept;

Where OP is: unary +, unary -.

Constructs and returns a new instance of the SYCL range class template with the same dimensionality as the rhs SYCL range, where each element of the new SYCL range instance is the result of an element-wise OP operator on the rhs SYCL range.

range& operatorOP(range& rhs) noexcept;

Where OP is: prefix ++, prefix --.

Assigns each element of the rhs range instance with the result of an element-wise OP operator on each element of the rhs range and returns this range.

range operatorOP(range& lhs, int) noexcept;

Where OP is: postfix ++, postfix --.

Make a copy of the lhs range. Assigns each element of the lhs range instance with the result of an element-wise OP operator on each element of the lhs range. Then return the initial copy of the range.

4.9.1.2. nd_range class
 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
namespace sycl {
template <int Dimensions = 1> class nd_range {
 public:
  static constexpr int dimensions = Dimensions;

  /* -- common interface members -- */

  // The offset is deprecated in SYCL 2020.
  nd_range(range<Dimensions> globalSize, range<Dimensions> localSize,
           id<Dimensions> offset = id<Dimensions>()) noexcept;

  range<Dimensions> get_global_range() const noexcept;
  range<Dimensions> get_local_range() const noexcept;
  range<Dimensions> get_group_range() const noexcept;
  id<Dimensions> get_offset() const noexcept; // Deprecated in SYCL 2020.
};
} // namespace sycl

nd_range<int Dimensions> defines the iteration domain of both the work-groups and the overall dispatch. To define this the nd_range comprises two ranges: the whole range over which the kernel is to be executed, and the range of each work group.

The SYCL nd_range class template provides the common by-value semantics (see Section 4.5.3).

A synopsis of the SYCL nd_range class is provided below. The constructors and member functions of the SYCL nd_range class are listed in Table 81 and Table 82 respectively. The additional common special member functions and common member functions are listed in Section 4.5.3 in Table 9 and Table 10 respectively.

Table 81. Constructors of the nd_range class
Constructor Description
nd_range<Dimensions>(
range<Dimensions> globalSize,
    range<Dimensions> localSize,
    id<Dimensions> offset = id<Dimensions>()) noexcept;

Construct an nd_range from the local and global constituent ranges. Supplying the option offset is deprecated in SYCL 2020. If the offset is not provided it will default to no offset.

Table 82. Member functions for the nd_range class
Member function Description
range<Dimensions> get_global_range() const noexcept;

Return the constituent global range.

range<Dimensions> get_local_range() const noexcept;

Return the constituent local range.

range<Dimensions> get_group_range() const noexcept;

Return a range representing the number of groups in each dimension. This range would result from globalSize/localSize as provided on construction.

id<Dimensions> get_offset() const noexcept;
    // Deprecated in SYCL 2020.

Deprecated in SYCL 2020. Return the constituent offset.

4.9.1.3. id class

id<int Dimensions> is a vector of Dimensions that is used to represent an id into a global or local range. It can be used as an index in an accessor of the same rank. The subscript operator (operator[](n)) returns the component n as a std::size_t.

The SYCL id class template provides the common by-value semantics (see Section 4.5.3).

A synopsis of the SYCL id class is provided below. The constructors, member functions and non-member functions of the SYCL id class are listed in Table 83, Table 84 and Table 85 respectively. The additional common special member functions and common member functions are listed in Section 4.5.3 in Table 9 and Table 10 respectively.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
namespace sycl {
template <int Dimensions = 1> class id {
 public:
  static constexpr int dimensions = Dimensions;

  id() noexcept;

  /* The following constructor is only available in the id class
   * specialization where: Dimensions==1 */
  id(std::size_t dim0) noexcept;
  /* The following constructor is only available in the id class
   * specialization where: Dimensions==2 */
  id(std::size_t dim0, std::size_t dim1) noexcept;
  /* The following constructor is only available in the id class
   * specialization where: Dimensions==3 */
  id(std::size_t dim0, std::size_t dim1, std::size_t dim2) noexcept;

  /* -- common interface members -- */

  id(const range<Dimensions>& range) noexcept;
  id(const item<Dimensions>& item) noexcept;

  std::size_t get(int dimension) const noexcept;
  std::size_t& operator[](int dimension) noexcept;
  std::size_t operator[](int dimension) const noexcept;

  // only available if Dimensions == 1
  operator std::size_t() const noexcept;

  // OP is: +, -, *, /, %, <<, >>, &, |, ^, &&, ||, <, >, <=, >=
  friend id operatorOP(const id& lhs, const id& rhs) noexcept { /* ... */
  }

  // OP is: +, -, *, /, %, <<, >>, &, |, ^, &&, ||, <, >, <=, >=
  // Available only when std::is_integral_v<T> is true
  template <typename T>
  friend id operatorOP(const id& lhs, const T& rhs) noexcept { /* ... */
  }

  // OP is: +, -, *, /, %, <<, >>, &, |, ^, &&, ||, <, >, <=, >=
  // Available only when std::is_integral_v<T> is true
  template <typename T>
  friend id operatorOP(const T& lhs, const id& rhs) noexcept { /* ... */
  }

  // OP is: +=, -=, *=, /=, %=, <<=, >>=, &=, |=, ^=
  friend id& operatorOP(id& lhs, const id& rhs) noexcept { /* ... */
  }

  // OP is: +=, -=, *=, /=, %=, <<=, >>=, &=, |=, ^=
  // Available only when std::is_integral_v<T> is true
  template <typename T>
  friend id& operatorOP(id& lhs, const T& rhs) noexcept { /* ... */
  }

  // OP is unary +, -
  friend id operatorOP(const id& rhs) noexcept { /* ... */
  }

  // OP is prefix ++, --
  friend id& operatorOP(id& rhs) noexcept { /* ... */
  }

  // OP is postfix ++, --
  friend id operatorOP(id& lhs, int) noexcept { /* ... */
  }
};

// Deduction guides
id(std::size_t)->id<1>;
id(std::size_t, std::size_t)->id<2>;
id(std::size_t, std::size_t, std::size_t)->id<3>;

} // namespace sycl
Table 83. Constructors of the id class template
Constructor Description
id() noexcept;

Construct a SYCL id with the value 0 for each dimension.

id(std::size_t dim0) noexcept;

Construct a 1D id with value dim0. Only valid when the template parameter Dimensions is equal to 1.

id(std::size_t dim0, std::size_t dim1) noexcept;

Construct a 2D id with values dim0, dim1. Only valid when the template parameter Dimensions is equal to 2.

id(std::size_t dim0, std::size_t dim1, std::size_t dim2) noexcept;

Construct a 3D id with values dim0, dim1, dim2. Only valid when the template parameter Dimensions is equal to 3.

id(const range<Dimensions>& range) noexcept;

Construct an id from the dimensions of range.

id(const item<Dimensions>& item) noexcept;

Construct an id from item.get_id().

Table 84. Member functions of the id class template
Member function Description
std::size_t get(int dimension) const noexcept;

Return the value of the requested dimension of this id object. Results in undefined behavior if dimension is not in the range [0, Dimensions).

std::size_t& operator[](int dimension) noexcept;

Return a reference to the requested dimension of the id object. Results in undefined behavior if dimension is not in the range [0, Dimensions).

std::size_t operator[](int dimension) const noexcept;

Return the value of the requested dimension of the id object. Results in undefined behavior if dimension is not in the range [0, Dimensions).

operator std::size_t() const noexcept;

Available only when: Dimensions == 1

Returns the same value as get(0).

Table 85. Hidden friend functions of the id class template
Hidden friend function Description
id operatorOP(const id& lhs, const id& rhs) noexcept;

Where OP is: +, -, *, /, %, <<, >>, &, |, ^, &&, ||, <, >, <=, >=.

Constructs and returns a new instance of the SYCL id class template with the same dimensionality as lhs id, where each element of the new SYCL id instance is the result of an element-wise OP operator between each element of lhs id and each element of the rhs id. If the operator returns a bool the result is then cast to std::size_t.

template <typename T>
id operatorOP(const id& lhs, const T& rhs) noexcept;

Where OP is: +, -, *, /, %, <<, >>, &, |, ^, &&, ||, <, >, <=, >=.

Constructs and returns a new instance of the SYCL id class template with the same dimensionality as lhs id, where each element of the new SYCL id instance is the result of an element-wise OP operator between each element of lhs id and the rhs integral type. If the operator returns a bool the result is then cast to std::size_t.

template <typename T>
id operatorOP(const T& lhs, const id& rhs) noexcept;

Where OP is: +, -, *, /, %, <<, >>, &, |, ^, &&, ||, <, >, <=, >=.

Constructs and returns a new instance of the SYCL id class template with the same dimensionality as the rhs SYCL id, where each element of the new SYCL id instance is the result of an element-wise OP operator between the lhs integral type and each element of the rhs SYCL id. If the operator returns a bool the result is then cast to std::size_t.

id& operatorOP(id& lhs, const id& rhs) noexcept;

Where OP is: +=, -=,*=, /=, %=, <<=, >>=, &=, |=, ^=.

Assigns each element of lhs id instance with the result of an element-wise OP operator between each element of lhs id and each element of the rhs id and returns lhs id. If the operator returns a bool the result is then cast to std::size_t.

template <typename T>
id& operatorOP(id& lhs, const T& rhs) noexcept;

Where OP is: +=, -=,*=, /=, %=, <<=, >>=, &=, |=, ^=.

Assigns each element of lhs id instance with the result of an element-wise OP operator between each element of lhs id and the rhs integral type and returns lhs id. If the operator returns a bool the result is then cast to std::size_t.

id operatorOP(const id& rhs) noexcept;

Where OP is: unary +, unary -.

Constructs and returns a new instance of the SYCL id class template with the same dimensionality as the rhs SYCL id, where each element of the new SYCL id instance is the result of an element-wise OP operator on the rhs SYCL id.

id& operatorOP(id& rhs) noexcept;

Where OP is: prefix ++, prefix --.

Assigns each element of the rhs id instance with the result of an element-wise OP operator on each element of the rhs id and returns this id.

id operatorOP(id& lhs, int) noexcept;

Where OP is: postfix ++, postfix --.

Make a copy of the lhs id. Assigns each element of the lhs id instance with the result of an element-wise OP operator on each element of the lhs id. Then return the initial copy of the id.

4.9.1.4. item class

The item class template identifies an instance of a kernel function object executing at each point in a range. It encapsulates enough information to identify the work-item’s range of possible values and its ID in that range.

The implementation constructs instances of the item class template. Applications that attempt to default construct an item object are ill formed, and the implementation must issue a diagnostic in this case.

[Note: If using C++17, implementations must ensure that the item class template is not an aggregate. — end note]

The item class template can optionally carry the offset of the range if provided to the parallel_for; note this is deprecated in SYCL 2020.

The SYCL item class template provides the common by-value semantics (see Section 4.5.3).

A synopsis of the SYCL item class is provided below. The member functions of the SYCL item class are listed in Table 84. The additional common special member functions and common member functions are listed in Section 4.5.3 in Table 9 and Table 10 respectively.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
namespace sycl {
template <int Dimensions = 1, bool WithOffset = true> class item {
 public:
  static constexpr int dimensions = Dimensions;

  item() = delete;

  /* -- common interface members -- */

  id<Dimensions> get_id() const noexcept;

  std::size_t get_id(int dimension) const noexcept;

  std::size_t operator[](int dimension) const noexcept;

  range<Dimensions> get_range() const noexcept;

  std::size_t get_range(int dimension) const noexcept;

  // Deprecated in SYCL 2020.
  // only available if WithOffset is true
  id<Dimensions> get_offset() const noexcept;

  // Deprecated in SYCL 2020.
  // only available if WithOffset is false
  operator item<Dimensions, true>() const noexcept;

  // only available if Dimensions == 1
  operator std::size_t() const noexcept;

  std::size_t get_linear_id() const noexcept;
};
} // namespace sycl
Table 86. Member functions for the item class
Member function Description
id<Dimensions> get_id() const noexcept;

Return the constituent id representing the work-item’s position in the iteration space.

std::size_t get_id(int dimension) const noexcept;

Equivalent to return get_id()[dimension].

std::size_t operator[](int dimension) const noexcept;

Equivalent to return get_id(dimension).

range<Dimensions> get_range() const noexcept;

Returns a range representing the dimensions of the range of possible values of the item.

std::size_t get_range(int dimension) const noexcept;

Equivalent to return get_range().get(dimension).

id<Dimensions> get_offset() const noexcept;
    // Deprecated in SYCL 2020.

Deprecated in SYCL 2020. Returns an id representing the n-dimensional offset provided to the parallel_for and that is added by the runtime to the global-ID of each work-item, if this item represents a global range. For an item converted from an item with no offset this will always return an id of all 0 values.

This member function is only available if WithOffset is true.

operator item<Dimensions, true>() const noexcept;
    // Deprecated in SYCL 2020.

Deprecated in SYCL 2020.

Available only when: WithOffset == false

Returns an item representing the same information as the object holds but also includes the offset set to 0. This conversion allow users to seamlessly write code that assumes an offset and still provides an offset-less item.

operator std::size_t() const noexcept;

Available only when: Dimensions == 1

Returns the same value as get_id(0).

std::size_t get_linear_id() const noexcept;

Return the id as a linear index value. Calculating a linear address from the multi-dimensional index follows Section 3.11.1.

4.9.1.5. nd_item class

The nd_item class template identifies an instance of a kernel function object executing at each point in an nd-range. It encapsulates enough information to identify the work-item’s local and global ids, the work-group id and also provides access to the group and sub_group classes.

The implementation constructs instances of the nd_item class template. Applications that attempt to default construct an nd_item object are ill formed, and the implementation must issue a diagnostic in this case.

[Note: If using C++17, implementations must ensure that the nd_item class template is not an aggregate. — end note]

The SYCL nd_item class template provides the common by-value semantics (see Section 4.5.3).

A synopsis of the SYCL nd_item class is provided below. The member functions of the SYCL nd_item class are listed in Table 87. The additional common special member functions and common member functions are listed in Section 4.5.3 in Table 9 and Table 10 respectively.

  1
  2
  3
  4
  5
  6
  7
  8
  9
 10
 11
 12
 13
 14
 15
 16
 17
 18
 19
 20
 21
 22
 23
 24
 25
 26
 27
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
namespace sycl {
template <int Dimensions = 1> class nd_item {
 public:
  static constexpr int dimensions = Dimensions;

  nd_item() = delete;

  /* -- common interface members -- */

  id<Dimensions> get_global_id() const noexcept;

  std::size_t get_global_id(int dimension) const noexcept;

  std::size_t get_global_linear_id() const noexcept;

  id<Dimensions> get_local_id() const noexcept;

  std::size_t get_local_id(int dimension) const noexcept;

  std::size_t get_local_linear_id() const noexcept;

  group<Dimensions> get_group() const noexcept;

  sub_group get_sub_group() const noexcept;

  std::size_t get_group(int dimension) const noexcept;

  std::size_t get_group_linear_id() const noexcept;

  range<Dimensions> get_group_range() const noexcept;

  std::size_t get_group_range(int dimension) const noexcept;

  range<Dimensions> get_global_range() const noexcept;

  std::size_t get_global_range(int dimension) const noexcept;

  range<Dimensions> get_local_range() const noexcept;

  std::size_t get_local_range(int dimension) const noexcept;

  // Deprecated in SYCL 2020.
  id<Dimensions> get_offset() const noexcept;

  nd_range<Dimensions> get_nd_range() const noexcept;

  // Deprecated in SYCL 2020. 
  template <typename DataT>
  device_event async_work_group_copy(local_ptr<DataT> dest,
                                     global_ptr<DataT> src,
                                     std::size_t numElements) const noexcept;

  // Deprecated in SYCL 2020.
  template <typename DataT>
  device_event async_work_group_copy(global_ptr<DataT> dest,
                                     local_ptr<DataT> src,
                                     std::size_t numElements) const noexcept;

  // Deprecated in SYCL 2020.
  template <typename DataT>
  device_event async_work_group_copy(local_ptr<DataT> dest,
                                     global_ptr<DataT> src,
                                     std::size_t numElements,
                                     std::size_t srcStride) const noexcept;

  // Deprecated in SYCL 2020.
  template <typename DataT>
  device_event async_work_group_copy(global_ptr<DataT> dest,
                                     local_ptr<DataT> src,
                                     std::size_t numElements,
                                     std::size_t destStride) const noexcept;

  /* Available only when: (std::is_same_v<DestDataT,
       std::remove_const_t<SrcDataT>> == true) */
  template <typename DestDataT, typename SrcDataT>
  device_event async_work_group_copy(decorated_local_ptr<DestDataT> dest,
                                     decorated_global_ptr<SrcDataT> src,
                                     std::size_t numElements) const noexcept;

  /* Available only when: (std::is_same_v<DestDataT,
       std::remove_const_t<SrcDataT>> == true) */
  template <typename DestDataT, typename SrcDataT>
  device_event async_work_group_copy(decorated_global_ptr<DestDataT> dest,
                                     decorated_local_ptr<SrcDataT> src,
                                     std::size_t numElements) const noexcept;

  /* Available only when: (std::is_same_v<DestDataT,
       std::remove_const_t<SrcDataT>> == true) */
  template <typename DestDataT, typename SrcDataT>
  device_event async_work_group_copy(decorated_local_ptr<DestDataT> dest,
                                     decorated_global_ptr<SrcDataT> src,
                                     std::size_t numElements,
                                     std::size_t srcStride) const noexcept;

  /* Available only when: (std::is_same_v<DestDataT,
       std::remove_const_t<SrcDataT>> == true) */
  template <typename DestDataT, typename SrcDataT>
  device_event async_work_group_copy(decorated_global_ptr<DestDataT> dest,
                                     decorated_local_ptr<SrcDataT> src,
                                     std::size_t numElements,
                                     std::size_t destStride) const noexcept;

  template <typename... EventTN> void wait_for(EventTN... events) const noexcept;
};
} // namespace sycl
Table 87. Member functions for the nd_item class
Member function Description
id<Dimensions> get_global_id() const noexcept;

Return the constituent global id representing the work-item’s position in the global iteration space.

std::size_t get_global_id(int dimension) const noexcept;

Return the constituent element of the global id representing the work-item’s position in the nd-range in the given Dimension. Results in undefined behavior if dimension is not in the range [0, Dimensions).

std::size_t get_global_linear_id() const noexcept;

Return the constituent global id as a linear index value, representing the work-item’s position in the global iteration space. The linear address is calculated from the multi-dimensional index by first subtracting the offset and then following Section 3.11.1.

id<Dimensions> get_local_id() const noexcept;

Return the constituent local id representing the work-item’s position within the current work-group.

std::size_t get_local_id(int dimension) const noexcept;

Return the constituent element of the local id representing the work-item’s position within the current work-group in the given Dimension. Results in undefined behavior if dimension is not in the range [0, Dimensions).

std::size_t get_local_linear_id() const noexcept;

Return the constituent local id as a linear index value, representing the work-item’s position within the current work-group. The linear address is calculated from the multi-dimensional index following Section 3.11.1.

group<Dimensions> get_group() const noexcept;

Return the constituent work-group, group representing the work-group's position within the overall nd-range.

sub_group get_sub_group() const noexcept;

Return a sub_group representing the sub-group to which the work-item belongs.

std::size_t get_group(int dimension) const noexcept;

Return the constituent element of the group id representing the work-group’s position within the overall nd_range in the given Dimension. Results in undefined behavior if dimension is not in the range [0, Dimensions).

std::size_t get_group_linear_id() const noexcept;

Return the group id as a linear index value. Calculating a linear address from a multi-dimensional index follows Section 3.11.1.

range<Dimensions> get_group_range() const noexcept;

Returns the number of work-groups in the iteration space.

std::size_t get_group_range(int dimension) const noexcept;

Return the number of work-groups for Dimension in the iteration space. Results in undefined behavior if dimension is not in the range [0, Dimensions).

range<Dimensions> get_global_range() const noexcept;

Returns a range representing the dimensions of the global iteration space.

std::size_t get_global_range(int dimension) const noexcept;

Equivalent to return get_global_range().get(dimension).

range<Dimensions> get_local_range() const noexcept;

Returns a range representing the dimensions of the current work-group.

std::size_t get_local_range(int dimension) const noexcept;

Equivalent to return get_local_range().get(dimension).

id<Dimensions> get_offset() const noexcept;
    // Deprecated in SYCL 2020.

Deprecated in SYCL 2020. Returns an id representing the n-dimensional offset provided to the constructor of the nd_range and that is added by the runtime to the global id of each work-item.

nd_range<Dimensions> get_nd_range() const noexcept;

Returns the nd_range of the current execution.

template <typename DataT>
device_event async_work_group_copy(local_ptr<DataT> dest,
                                   global_ptr<DataT> src,
                                   std::size_t numElements) const noexcept;

Deprecated in SYCL 2020. Has the same effect as the overload taking decorated_local_ptr and decorated_global_ptr except that the dest and src parameters are multi_ptr with access::decorated::legacy.

template <typename DataT>
device_event async_work_group_copy(global_ptr<DataT> dest,
                                   local_ptr<DataT> src,
                                   std::size_t numElements) const noexcept;

Deprecated in SYCL 2020. Has the same effect as the overload taking decorated_local_ptr and decorated_global_ptr except that the dest and src parameters are multi_ptr with access::decorated::legacy.

template <typename DataT>
device_event async_work_group_copy(local_ptr<DataT> dest,
                                   global_ptr<DataT> src,
                                   std::size_t numElements, std::size_t srcStride) const noexcept;

Deprecated in SYCL 2020. Has the same effect as the overload taking decorated_local_ptr and decorated_global_ptr except that the dest and src parameters are multi_ptr with access::decorated::legacy.

template <typename DataT>
device_event async_work_group_copy(global_ptr<DataT> dest,
                                   local_ptr<DataT> src,
                                   std::size_t numElements, std::size_t destStride) const noexcept;

Deprecated in SYCL 2020. Has the same effect as the overload taking decorated_local_ptr and decorated_global_ptr except that the dest and src parameters are multi_ptr with access::decorated::legacy.

template <typename DestDataT, typename SrcDataT>
device_event async_work_group_copy(decorated_local_ptr<DestDataT> dest,
                                   decorated_global_ptr<SrcDataT> src,
                                   std::size_t numElements) const noexcept;

Available only when: (std::is_same_v<DestDataT, std::remove_const_t<SrcDataT>> == true)

Permitted types for DataT are all scalar and vector types. Asynchronously copies a number of elements specified by numElements from the source pointer src to destination pointer dest and returns a SYCL device_event which can be used to wait on the completion of the copy.

template <typename DestDataT, typename SrcDataT>
device_event async_work_group_copy(decorated_global_ptr<DestDataT> dest,
                                   decorated_local_ptr<SrcDataT> src,
                                   std::size_t numElements) const noexcept;

Available only when: (std::is_same_v<DestDataT, std::remove_const_t<SrcDataT>> == true)

Permitted types for DataT are all scalar and vector types. Asynchronously copies a number of elements specified by numElements from the source pointer src to destination pointer dest and returns a SYCL device_event which can be used to wait on the completion of the copy.

template <typename DestDataT, typename SrcDataT>
device_event async_work_group_copy(decorated_local_ptr<DestDataT> dest,
                                   decorated_global_ptr<SrcDataT> src,
                                   std::size_t numElements, std::size_t srcStride) const noexcept;

Available only when: (std::is_same_v<DestDataT, std::remove_const_t<SrcDataT>> == true)

Permitted types for DataT are all scalar and vector types. Asynchronously copies a number of elements specified by numElements from the source pointer src to destination pointer dest with a source stride specified by srcStride and returns a SYCL device_event which can be used to wait on the completion of the copy.

template <typename DestDataT, SrcDataT>
device_event async_work_group_copy(decorated_global_ptr<DestDataT> dest,
                                   decorated_local_ptr<SrcDataT> src,
                                   std::size_t numElements, std::size_t destStride) const noexcept;

Available only when: (std::is_same_v<DestDataT, std::remove_const_t<SrcDataT>> == true)

Permitted types for DataT are all scalar and vector types. Asynchronously copies a number of elements specified by numElements from the source pointer src to destination pointer dest with a destination stride specified by destStride and returns a SYCL device_event which can be used to wait on the completion of the copy.

template <typename... EventTN> void wait_for(EventTN... events) const noexcept;

Permitted type for EventTN is device_event. Waits for the asynchronous operations associated with each device_event to complete.

4.9.1.6. h_item class (deprecated)

The h_item class is deprecated in SYCL 2020.

The h_item class template identifies an instance of a group::parallel_for_work_item function object executing at each point in a local range<int Dimensions> passed to a parallel_for_work_item call or to the corresponding parallel_for_work_group call if no range is passed to the parallel_for_work_item call. It encapsulates enough information to identify the work-item's local and global items according to the information given to parallel_for_work_group (physical ids) as well as the work-item's logical local items in the logical local range. All returned items objects are offset-less.

The implementation constructs instances of the h_item class template. Applications that attempt to default construct an h_item object are ill formed, and the implementation must issue a diagnostic in this case.

[Note: If using C++17, implementations must ensure that the h_item class template is not an aggregate. — end note]

The SYCL h_item class template provides the common by-value semantics (see Section 4.5.3).

A synopsis of the SYCL h_item class is provided below. The member functions of the SYCL h_item class are listed in Table 88. The additional common special member functions and common member functions are listed in Section 4.5.3 in Table 9 and Table 10 respectively.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
namespace sycl {
/* Deprecated in SYCL 2020 */
template <int Dimensions> class h_item {
 public:
  static constexpr int dimensions = Dimensions;

  h_item() = delete;

  /* -- common interface members -- */

  item<Dimensions, false> get_global() const noexcept;

  item<Dimensions, false> get_local() const noexcept;

  item<Dimensions, false> get_logical_local() const noexcept;

  item<Dimensions, false> get_physical_local() const noexcept;

  range<Dimensions> get_global_range() const noexcept;

  std::size_t get_global_range(int dimension) const noexcept;

  id<Dimensions> get_global_id() const noexcept;

  std::size_t get_global_id(int dimension) const noexcept;

  range<Dimensions> get_local_range() const noexcept;

  std::size_t get_local_range(int dimension) const noexcept;

  id<Dimensions> get_local_id() const noexcept;

  std::size_t get_local_id(int dimension) const noexcept;

  range<Dimensions> get_logical_local_range() const noexcept;

  std::size_t get_logical_local_range(int dimension) const noexcept;

  id<Dimensions> get_logical_local_id() const noexcept;

  std::size_t get_logical_local_id(int dimension) const noexcept;

  range<Dimensions> get_physical_local_range() const noexcept;

  std::size_t get_physical_local_range(int dimension) const noexcept;

  id<Dimensions> get_physical_local_id() const noexcept;

  std::size_t get_physical_local_id(int dimension) const noexcept;
};
} // namespace sycl
Table 88. Member functions for the h_item class
Member function Description
item<Dimensions, false> get_global() const noexcept;

Return the constituent global item representing the work-item’s position in the global iteration space as provided upon kernel invocation.

item<Dimensions, false> get_local() const noexcept;

Return the same value as get_logical_local().

item<Dimensions, false> get_logical_local() const noexcept;

Return the constituent element of the logical local item work-item’s position in the local iteration space as provided upon the invocation of the group::parallel_for_work_item.

If the group::parallel_for_work_item was called without any logical local range then the member function returns the physical local item.

A physical id can be computed from a logical id by getting the remainder of the integer division of the logical id and the physical range: get_logical_local().get() % get_physical_local.get_range() == get_physical_local().get().

item<Dimensions, false> get_physical_local() const noexcept;

Return the constituent element of the physical local item work-item’s position in the local iteration space as provided (by the user or the runtime) upon the kernel invocation.

range<Dimensions> get_global_range() const noexcept;

Equivalent to return get_global().get_range()

std::size_t get_global_range(int dimension) const noexcept;

Equivalent to return get_global().get_range(dimension)

id<Dimensions> get_global_id() const noexcept;

Equivalent to return get_global().get_id()

std::size_t get_global_id(int dimension) const noexcept;

Equivalent to return get_global().get_id(dimension)

range<Dimensions> get_local_range() const noexcept;

Equivalent to return get_local().get_range()

std::size_t get_local_range(int dimension) const noexcept;

Equivalent to return get_local().get_range(dimension)

id<Dimensions> get_local_id() const noexcept;

Equivalent to return get_local().get_id()

std::size_t get_local_id(int dimension) const noexcept;

Equivalent to return get_local().get_id(dimension)

range<Dimensions> get_logical_local_range() const noexcept;

Equivalent to return get_logical_local().get_range()

std::size_t get_logical_local_range(int dimension) const noexcept;

Equivalent to return get_logical_local().get_range(dimension)

id<Dimensions> get_logical_local_id() const noexcept;

Equivalent to return get_logical_local().get_id()

std::size_t get_logical_local_id(int dimension) const noexcept;

Equivalent to return get_logical_local().get_id(dimension)

range<Dimensions> get_physical_local_range() const noexcept;

Equivalent to return get_physical_local().get_range()

std::size_t get_physical_local_range(int dimension) const noexcept;

Equivalent to return get_physical_local().get_range(dimension)

id<Dimensions> get_physical_local_id() const noexcept;

Equivalent to return get_physical_local().get_id()

std::size_t get_physical_local_id(int dimension) const noexcept;

Equivalent to return get_physical_local().get_id(dimension)

4.9.1.7. group class

The group class template encapsulates all functionality required to represent a specific work-group within a kernel.

The set of work-items represented by an instance of the group class template is determined by the implementation, and there is subsequently no way for a user to construct arbitrary instances of the group class template. Applications that attempt to default construct a group object are ill formed, and the implementation must issue a diagnostic in this case.

[Note: If using C++17, implementations must ensure that the group class template is not an aggregate. — end note]

The local range stored in the group class is provided either by the programmer, when it is passed as an optional parameter to parallel_for_work_group, or by the runtime system when it selects the optimal work-group size. This allows the developer to always know how many work-items are in each executing work-group, even through the abstracted iteration range of the parallel_for_work_item loops.

The SYCL group class template provides the common by-value semantics (see Section 4.5.3).

A synopsis of the SYCL group class is provided below. The member functions of the SYCL group class are listed in Table 89. The additional common special member functions and common member functions are listed in Section 4.5.3 in Table 9 and Table 10 respectively.

  1
  2
  3
  4
  5
  6
  7
  8
  9
 10
 11
 12
 13
 14
 15
 16
 17
 18
 19
 20
 21
 22
 23
 24
 25
 26
 27
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
namespace sycl {
template <int Dimensions = 1> class group {
 public:
  using id_type = id<Dimensions>;
  using range_type = range<Dimensions>;
  using linear_id_type = std::size_t;
  static constexpr int dimensions = Dimensions;
  static constexpr memory_scope fence_scope = memory_scope::work_group;

  group() = delete;

  /* -- common interface members -- */

  id<Dimensions> get_group_id() const noexcept;

  std::size_t get_group_id(int dimension) const noexcept;

  id<Dimensions> get_local_id() const noexcept;

  std::size_t get_local_id(int dimension) const noexcept;

  range<Dimensions> get_local_range() const noexcept;

  std::size_t get_local_range(int dimension) const noexcept;

  range<Dimensions> get_group_range() const noexcept;

  std::size_t get_group_range(int dimension) const noexcept;

  range<Dimensions> get_max_local_range() const noexcept;

  std::size_t operator[](int dimension) const noexcept;

  std::size_t get_group_linear_id() const noexcept;

  std::size_t get_local_linear_id() const noexcept;

  std::size_t get_group_linear_range() const noexcept;

  std::size_t get_local_linear_range() const noexcept;

  bool leader() const noexcept;

  // Deprecated in SYCL 2020.
  template <typename WorkItemFunctionT>
  void parallel_for_work_item(const WorkItemFunctionT& func) const noexcept;

  // Deprecated in SYCL 2020.
  template <typename WorkItemFunctionT>
  void parallel_for_work_item(range<Dimensions> logicalRange,
                              const WorkItemFunctionT& func) const noexcept;

  // Deprecated in SYCL 2020. 
  template <typename DataT>
  device_event async_work_group_copy(local_ptr<DataT> dest,
                                     global_ptr<DataT> src,
                                     std::size_t numElements) const noexcept;

  // Deprecated in SYCL 2020.
  template <typename DataT>
  device_event async_work_group_copy(global_ptr<DataT> dest,
                                     local_ptr<DataT> src,
                                     std::size_t numElements) const noexcept;

  // Deprecated in SYCL 2020.
  template <typename DataT>
  device_event async_work_group_copy(local_ptr<DataT> dest,
                                     global_ptr<DataT> src,
                                     std::size_t numElements,
                                     std::size_t srcStride) const noexcept;

  // Deprecated in SYCL 2020.
  template <typename DataT>
  device_event async_work_group_copy(global_ptr<DataT> dest,
                                     local_ptr<DataT> src,
                                     std::size_t numElements,
                                     std::size_t destStride) const noexcept;

  /* Available only when: (std::is_same_v<DestDataT,
       std::remove_const_t<SrcDataT>> == true) */
  template <typename DestDataT, typename SrcDataT>
  device_event async_work_group_copy(decorated_local_ptr<DestDataT> dest,
                                     decorated_global_ptr<SrcDataT> src,
                                     std::size_t numElements) const noexcept;

  /* Available only when: (std::is_same_v<DestDataT,
       std::remove_const_t<SrcDataT>> == true) */
  template <typename DestDataT, typename SrcDataT>
  device_event async_work_group_copy(decorated_global_ptr<DestDataT> dest,
                                     decorated_local_ptr<SrcDataT> src,
                                     std::size_t numElements) const noexcept;

  /* Available only when: (std::is_same_v<DestDataT,
       std::remove_const_t<SrcDataT>> == true) */
  template <typename DestDataT, typename SrcDataT>
  device_event async_work_group_copy(decorated_local_ptr<DestDataT> dest,
                                     decorated_global_ptr<SrcDataT> src,
                                     std::size_t numElements,
                                     std::size_t srcStride) const noexcept;

  /* Available only when: (std::is_same_v<DestDataT,
       std::remove_const_t<SrcDataT>> == true) */
  template <typename DestDataT, typename SrcDataT>
  device_event async_work_group_copy(decorated_global_ptr<DestDataT> dest,
                                     decorated_local_ptr<SrcDataT> src,
                                     std::size_t numElements,
                                     std::size_t destStride) const noexcept;

  template <typename... EventTN> void wait_for(EventTN... events) const noexcept;
};
} // namespace sycl
Table 89. Member functions for the group class
Member function Description
id<Dimensions> get_group_id() const noexcept;

Return an id representing the index of the work-group within the global nd-range for every dimension. Since the work-items in a work-group have a defined position within the global nd-range, the returned group id can be used along with the local id to uniquely identify the work-item in the global nd-range.

std::size_t get_group_id(int dimension) const noexcept;

Equivalent to return get_group_id()[dimension].

id<Dimensions> get_local_id() const noexcept;

Return a SYCL id representing the calling work-item’s position within the work-group.

It is undefined behavior for this member function to be invoked from within a parallel_for_work_item context.

std::size_t get_local_id(int dimension) const noexcept;

Equivalent to return get_local_id()[dimension].

It is undefined behavior for this member function to be invoked from within a parallel_for_work_item context.

range<Dimensions> get_local_range() const noexcept;

Return a SYCL range representing all dimensions of the local range. This local range may have been provided by the programmer, or chosen by the SYCL runtime.

std::size_t get_local_range(int dimension) const noexcept;

Equivalent to return get_local_range()[dimension].

range<Dimensions> get_group_range() const noexcept;

Return a range representing the number of work-groups in the nd_range.

std::size_t get_group_range(int dimension) const noexcept;

Equivalent to return get_group_range()[dimension].

std::size_t operator[](int dimension) const noexecpt;

Equivalent to return get_group_id(dimension).

range<Dimensions> get_max_local_range() const noexcept;

Return a range representing the maximum number of work-items in any work-group in the nd_range.

std::size_t get_group_linear_id() const noexcept;

Get a linearized version of the work-group id. Calculating a linear work-group id from a multi-dimensional index follows Section 3.11.1.

std::size_t get_group_linear_range() const noexcept;

Return the total number of work-groups in the nd_range.

std::size_t get_local_linear_id() const noexcept;

Get a linearized version of the calling work-item’s local id. Calculating a linear local id from a multi-dimensional index follows Section 3.11.1.

It is undefined behavior for this member function to be invoked from within a parallel_for_work_item context.

std::size_t get_local_linear_range() const noexcept;

Return the total number of work-items in the work-group.

bool leader() const noexcept;

Return true for exactly one work-item in the work-group, if the calling work-item is the leader of the work-group, and false for all other work-items in the work-group.

The leader of the work-group is determined during construction of the work-group, and is invariant for the lifetime of the work-group. The leader of the work-group is guaranteed to be the work-item with a local id of 0.

template <typename WorkItemFunctionT>
void parallel_for_work_item(const WorkItemFunctionT& func) const noexcept;

Deprecated in SYCL 2020. Launch the work-items for this work-group.

func is a function object type with a public member function void F::operator()(h_item<Dimensions>) representing the work-item computation.

This member function can only be invoked within a parallel_for_work_group context. It is undefined behavior for this member function to be invoked from within the parallel_for_work_group form that does not define work-group size, because then the number of work-items that should execute the code is not defined. It is expected that this form of parallel_for_work_item is invoked within the parallel_for_work_group form that specifies the size of a work-group.

template <typename WorkItemFunctionT>
void parallel_for_work_item(range<Dimensions> logicalRange,
                            const WorkItemFunctionT& func) const noexcept;

Deprecated in SYCL 2020. Launch the work-items for this work-group using a logical local range. The function object func is executed as if the kernel were invoked with logicalRange as the local range. This new local range is emulated and may not map one-to-one with the physical range.

logicalRange is the new local range to be used. This range can be smaller or larger than the one used to invoke the kernel. func is a function object type with a public member function void F::operator()(h_item<Dimensions>) representing the work-item computation.

Note that the logical range does not need to be uniform across all work-groups in a kernel. For example the logical range may depend on a work-group varying query (e.g. group::get_linear_id), such that different work-groups in the same kernel invocation execute different logical range sizes.

This member function can only be invoked within a parallel_for_work_group context.

template <typename DataT>
device_event async_work_group_copy(local_ptr<DataT> dest,
                                   global_ptr<DataT> src,
                                   std::size_t numElements) const noexcept;

Deprecated in SYCL 2020. Has the same effect as the overload taking decorated_local_ptr and decorated_global_ptr except that the dest and src parameters are multi_ptr with access::decorated::legacy.

template <typename DataT>
device_event async_work_group_copy(global_ptr<DataT> dest,
                                   local_ptr<DataT> src,
                                   std::size_t numElements) const noexcept;

Deprecated in SYCL 2020. Has the same effect as the overload taking decorated_local_ptr and decorated_global_ptr except that the dest and src parameters are multi_ptr with access::decorated::legacy.

template <typename DataT>
device_event async_work_group_copy(local_ptr<DataT> dest,
                                   global_ptr<DataT> src,
                                   std::size_t numElements, std::size_t srcStride) const noexcept;

Deprecated in SYCL 2020. Has the same effect as the overload taking decorated_local_ptr and decorated_global_ptr except that the dest and src parameters are multi_ptr with access::decorated::legacy.

template <typename DataT>
device_event async_work_group_copy(global_ptr<DataT> dest,
                                   local_ptr<DataT> src,
                                   std::size_t numElements, std::size_t destStride) const noexcept;

Deprecated in SYCL 2020. Has the same effect as the overload taking decorated_local_ptr and decorated_global_ptr except that the dest and src parameters are multi_ptr with access::decorated::legacy.

template <typename DestDataT, typename SrcDataT>
device_event async_work_group_copy(decorated_global_ptr<DestDataT> dest,
                                   decorated_local_ptr<SrcDataT> src,
                                   std::size_t numElements) const noexcept;

Available only when: (std::is_same_v<DestDataT, std::remove_const_t<SrcDataT>> == true)

Permitted types for DataT are all scalar and vector types. Asynchronously copies a number of elements specified by numElements from the source pointer src to destination pointer dest and returns a SYCL device_event which can be used to wait on the completion of the copy.

template <typename DestDataT, typename SrcDataT>
device_event async_work_group_copy(decorated_local_ptr<DestDataT> dest,
                                   decorated_global_ptr<SrcDataT> src,
                                   std::size_t numElements, std::size_t srcStride) const noexcept;

Available only when: (std::is_same_v<DestDataT, std::remove_const_t<SrcDataT>> == true)

Permitted types for DataT are all scalar and vector types. Asynchronously copies a number of elements specified by numElements from the source pointer src to destination pointer dest with a source stride specified by srcStride and returns a SYCL device_event which can be used to wait on the completion of the copy.