ansaurus

Question

How to separate CUDA code into multiple files

Answer 1

+2 A:

You are including mykernel.cu in kernelsupport.cu, when you try to link the compiler sees mykernel.cu twice. You'll have to create a header defining TestDevice and include that instead.

re comment:

Something like this should work

// MyKernel.h
#ifndef mykernel_h
#define mykernel_h
__global__ void TestDevice(int* devicearray);
#endif

and then change the including file to

//KernelSupport.cu
#ifndef _KERNEL_SUPPORT_
#define _KERNEL_SUPPORT_

#include <iostream>
#include <MyKernel.h>
// ...

re your edit

As long as the header you use in c++ code doesn't have any cuda specific stuff (__kernel__,__global__, etc) you should be fine linking c++ and cuda code.

Scott Wales 2010-01-19 05:00:00

Please elaborate with a simple code example

Mr Bell 2010-01-19 05:38:47

Your MyKernel.h should have `void TestDeviceWrapper(dim3 grid, dim3 block, int *devicearray)` since when the KernelSupport.cu becomes KernelSupport.cpp cl.exe will not understand the __global__ syntax. Then in the MyKernel.cu, `TestDeviceWrapper()` just calls `TestDevice<<<>>>`.

Tom 2010-01-19 08:34:22

That sounds reasonable, the code given assumes it will be included in a cuda file, as is given in the question.

Scott Wales 2010-01-19 09:13:17

Yes, but he also says "The end result I am looking for here is to have a normal C++ application with something like Main.cpp with the int main() event and have things run from there." That was added in an edit to the question though.

Tom 2010-01-19 10:10:06

Answer 2

A:

The simple solution is to turn off building of your MyKernel.cu file.

Properties -> General -> Excluded from build

The better solution imo is to split your kernel into a cu and a cuh file, and include that, for example:

//kernel.cu
#include "kernel.cuh"
#include <cuda_runtime.h>

__global__ void increment_by_one_kernel(int* vals) {
  vals[threadIdx.x] += 1;
}

void increment_by_one(int* a) {
  int* a_d;

  cudaMalloc(&a_d, 1);
  cudaMemcpy(a_d, a, 1, cudaMemcpyHostToDevice);
  increment_by_one_kernel<<<1, 1>>>(a_d);
  cudaMemcpy(a, a_d, 1, cudaMemcpyDeviceToHost);

  cudaFree(a_d);
}

//kernel.cuh
#pragma once

void increment_by_one(int* a);

//main.cpp
#include "kernel.cuh"

int main() {
  int a[] = {1};

  increment_by_one(a);

  return 0;
}

thebaldwin 2010-01-19 05:00:55

Please elaborate with a simple code example

Mr Bell 2010-01-19 05:37:41

This will only work while you have your main in a .cu file. As soon as you put it into a .cpp file this is unsuitable.

Tom 2010-01-19 08:31:26

Once you split out all your CUDA/kernel code into appropriate cu/cuh files, there should be no problem renaming or moving your main to a cpp file. Please see my example, I'm unclear why it is unsuitable.

thebaldwin 2010-01-20 01:05:25

Answer 3

+1 A:

If you look at the CUDA SDK code examples, they have extern C defines that reference functions compiled from .cu files. This way, the .cu files are compiled by nvcc and only linked into the main program while the .cpp files are compiled normally.

For example, in marchingCubes_kernel.cu has the function body:

extern "C" void
launch_classifyVoxel( dim3 grid, dim3 threads, uint* voxelVerts, uint *voxelOccupied, uchar *volume,
                      uint3 gridSize, uint3 gridSizeShift, uint3 gridSizeMask, uint numVoxels,
                      float3 voxelSize, float isoValue)
{
    // calculate number of vertices need per voxel
    classifyVoxel<<<grid, threads>>>(voxelVerts, voxelOccupied, volume, 
                                     gridSize, gridSizeShift, gridSizeMask, 
                                     numVoxels, voxelSize, isoValue);
    cutilCheckMsg("classifyVoxel failed");
}

While in marchingCubes.cpp (where main() resides) just has a definition:

extern "C" void
launch_classifyVoxel( dim3 grid, dim3 threads, uint* voxelVerts, uint *voxelOccupied, uchar *volume,
                      uint3 gridSize, uint3 gridSizeShift, uint3 gridSizeMask, uint numVoxels,
                      float3 voxelSize, float isoValue);

You can put these in a .h file too.

tkerwin 2010-01-19 06:24:05

You should not need to use `extern "C"` in recent versions of the CUDA toolkit. In the past this was required since nvcc treated host code as C, however the default is now C++. Drop the `extern "C"`, it obfuscates the code!

Tom 2010-01-19 08:30:20

Good to know. They should update the SDK examples to reflect that. However, you still need to perform the CUDA call wrapping, I don't think there's any easy way around that.

tkerwin 2010-01-19 12:32:37

Yeah, the SDK samples haven't been updated since they were created, so while the newer ones reflect the latest standards the older ones are a little out-of-date. They do still illustrate the coding techniques though, if not the style.You are correct, there is no way to avoid the CUDA call wrapping. That makes total sense though, the triple chevron syntax (<<<>>>) is part of CUDA C and not C and hence you will need a CUDA C compiler (i.e. nvcc) to compile it. It's a small price to pay for the elegance of the Runtime API I think.

Tom 2010-01-19 15:11:54

Answer 4

+1 A:

Getting the separation is actually quite simple, please check out this answer for how to set it up. Then you simply put your host code in .cpp files and your device code in .cu files, the build rules tell Visual Studio how to link them together into the final executable.

The immediate problem in your code that you are defining the __global__ TestDevice function twice, once when you #include MyKernel.cu and once when you compile the MyKernel.cu independently.

You will need to put a wrapper into a .cu file too - at the moment you are calling TestDevice<<<>>> from your main function but when you move this into a .cpp file it will be compiled with cl.exe, which doesn't understand the <<<>>> syntax. Therefore you would simply call TestDeviceWrapper(griddim, blockdim, params) in the .cpp file and provide this function in your .cu file.

If you want an example, the SobolQRNG sample in the SDK achieves nice separation, although it still uses cutil and I would always recommend avoiding cutil.

Tom 2010-01-19 08:28:53

ansaurus

tags:

views:

answers:

How to separate CUDA code into multiple files

related questions