.NET bindings for transcribe.cpp: load GGUF speech-to-text models and transcribe audio (16 kHz mono float PCM) from C#.
Two ways to use it:
- A command-line tool β
dotnet tool install -g TranscribeCppSharp.Cligives youtranscribe, which downloads a curated model on first use and prints or exports the transcript. See Command-line tool. - A .NET library β the
TranscribeCppSharpwrapper (plus the native runtime package for your platform) for C# code. See Getting started.
π Full documentation β β the guide, the reference, and the measured limits with the method behind each one. Published from docs/; this README is the short version and the details are delegated to those pages.
The wrapper is platform-agnostic and ships no native binaries. A minimal install is the wrapper plus the native runtime package for your platform:
dotnet add package TranscribeCppSharp
dotnet add package TranscribeCppSharp.Native.linux-x64 # pick your platformIf you would rather not pick a platform (or want one package that works everywhere), install the bundle meta-package instead β it pulls the wrapper and every native runtime:
dotnet add package TranscribeCppSharp.BundleNative runtime packages: TranscribeCppSharp.Native.linux-x64,
.linux-arm64, .win-x64, .osx-arm64, .osx-x64.
Note: For Linux Alpine (musl) or other platforms, see Building from source. Like Using CUDA, a custom native build is picked up automatically when placed in the app output directory.
The wrapper resolves libtranscribe automatically in plain dotnet run scenarios (no
<RuntimeIdentifier> needed). If the native library is still missing at runtime (e.g. you
forgot the runtime package), the wrapper throws a DllNotFoundException that lists the
exact package to add for your platform and the paths it searched. It does not silently
produce a misleading error. See
Installation and
Native library loading.
The four snippets below are checked by a test: each is compared byte-for-byte against a
marked region in HighLevelApiTests.cs and samples/SmokeTest/Program.cs, so they
cannot drift from the API without failing the build
(ReadmeExamplesTest).
Model.Load initializes the compute backends automatically on first use, but you can (and for custom setups, should) do it explicitly with Backends.InitDefault():
Backends.InitDefault(); // optional: automatic in Model.Load, but explicit is clearer
var modelPath = TestConfig.ModelPath; // your GGUF model file, e.g. "test-models/ggml-tiny.bin"
var audioPath = TestConfig.AudioPath; // your WAV audio file, e.g. "test-audio/jfk.wav"
using var model = Model.Load(modelPath, p => p.WithBackend(BackendRequest.BackendCpu));
using var session = model.CreateSession();
var pcm = PcmExtensions.ReadWavToPcm(audioPath);
var transcript = session.Run(pcm);Omit the WithBackend call to get the default AUTO policy, which runs on the GPU
whenever one initializes β see Compute.
Backends.InitDefault(); // optional: automatic in Model.Load, but explicit is clearer
var modelPath = TestConfig.ModelPath; // your GGUF model file, e.g. "test-models/ggml-tiny.bin"
var audioPath = TestConfig.AudioPath; // your WAV audio file, e.g. "test-audio/jfk.wav"
using var model = Model.Load(modelPath, p => p.WithBackend(BackendRequest.BackendCpu));
using var session = model.CreateSession();
var pcm1 = PcmExtensions.ReadWavToPcm(audioPath);
var pcm2 = PcmExtensions.ReadWavToPcm(audioPath);
var results = Batch.Run(session, new[] { pcm1, pcm2 });stream.Begin();
int chunkSize = 16000; // 1 second
for (int i = 0; i < pcm.Length; i += chunkSize)
{
int length = Math.Min(chunkSize, pcm.Length - i);
var chunk = pcm.AsSpan(i, length);
stream.Feed(chunk);
}
stream.Complete();
var text = stream.GetCurrentText();A model is asked what it can do, not assumed. model.Supports(Feature.FeatureDiarization)
is the one worth calling before a run:
Backends.InitDefault(); // optional: automatic in Model.Load, but explicit is clearer
var modelPath = TestConfig.ModelPath; // your GGUF model file, e.g. "test-models/ggml-tiny.bin"
using var model = Model.Load(modelPath, p => p.WithBackend(BackendRequest.BackendCpu));
var supportsPnc = model.Supports(Feature.FeaturePnc);
var caps = model.GetCapabilities();Speaker attribution, long audio and windowing: Long audio and diarization.
Stated plainly, so nothing is implied. Each item has its own page with the detail.
- The library is not thread-safe. At most one run per model at a time, across all of
that model's sessions; parallel workers each need their own
Model. This is an upstream 0.x limitation, not one this wrapper imposes (Concurrency). - Streaming cannot attribute speakers.
transcribe_stream_paramshas nodiarizefield in transcribe.cpp v0.2.4, so the streaming API cannot request it; we do not invent one (Diarization). - No CUDA runtime ships in the packages. The bundled binaries are CPU + Vulkan on Windows/Linux and Metal on macOS; an NVIDIA GPU means supplying your own CUDA build (Using CUDA).
- The MIT license does not cover the models. They come from different ecosystems, some non-commercial, and this project neither bundles nor curates them (Model licenses).
- Speaker attribution is not verified in CI. It is covered only when the ~617 MB MOSS
asset is fetched with
WITH_DIARIZATION_MODEL=1. Until then, a regression that emptiedSpeakerSegmentswould go unnoticed (diarization). - There is no GPU in CI. The tool is tested end to end on the CPU, but GPU selection itself is covered by policy tests, not by a real run (what the CLI does not do).
The formats the decoder is verified to read are listed, with the tests that cover them, on Audio input.
This project is a packaging and binding effort only β the underlying library is not my work:
- The native library (
transcribe.cpp) is developed and owned by the transcribe.cpp authors (MIT License). - The bundled native components (ggml, etc.) are owned by their respective authors; their MIT license texts are distributed alongside the binaries in the
TranscribeCppSharp.Native.*packages. - I did not author the native library and claim no credit for it. This repository only adds:
- A C# interop layer (auto-generated P/Invoke bindings via
LibraryImport). - A high-level C# wrapper (
IDisposableresources, typed exceptions). - A command-line tool (
transcribe) and the model manifest it resolves against. - Pre-built native binaries packaged for .NET consumption.
- A C# interop layer (auto-generated P/Invoke bindings via
The transcribe.cpp project is an independent upstream project; bug reports about the native library itself should go to its repository.
# Download native libraries for your current platform
dotnet run --project tools/FetchNative
# Run unit and integration tests
./scripts/run-integration-tests.shPrerequisites, the opt-in diarization run, the checks that keep this documentation honest, and building from source: Development.
Security: report a vulnerability through the GitHub Security Advisory feature.
License: this project is licensed under the MIT License (matching transcribe.cpp).
The models you load are not covered by it β see
Model licenses.
Versioning: the wrapper and the native runtime are versioned separately on purpose;
TranscribeCppSharp 0.3.1 binds transcribe.cpp v0.2.4. See
Versioning & compatibility and
CHANGELOG.md.