Indexing & slicing¶
operator() does both element access and slicing, chosen by the argument types,
like NumPy's [].
Element access¶
All-integer arguments return a reference to one element (T&). Negative
indices count from the back:
t(1, 2, 3); // T& at (1,2,3)
t(0, -1); // row 0, last column
t(Int<1>(), j); // a static index folds; a runtime index is a value
t[1, 2, 3]; // C++23 only: multidimensional subscript, an exact alias of t(1,2,3)
operator[] is available when the compiler provides the C++23 multidimensional
subscript (__cpp_multidimensional_subscript) — it forwards to operator(), so
t[i, j], t[0, all, slice(1,4)] and t[ellipsis] all behave identically
(mdspan's own spelling). On C++17/20 use operator().
The negative-index wrap is a signed compare that folds away for static
(Int<k>()) and unsigned arguments. For runtime signed indices in a hot loop
where non-negative is guaranteed, compile with -DTNY_NO_NEGATIVE_INDEX to drop
the check.
at — one element as a rank-0 view¶
at(i...) takes the same all-integer indices but returns the element as a
rank-0 view instead of a T&, so the whole tensor API applies to a single
cell. Rank-0 tensors convert to and from T and have .item():
x.at(i,j) = 3; // write
float v = x.at(i,j); // read (implicit conversion to T)
float w = x.at(i,j).item(); // explicit read
x.at(i,j).atomic_add_(v); // atomic scatter into one cell
Unchecked accessors — uget / uat¶
The wrap that lets a negative index count from the back costs one compare-and-add
per axis. In a hot kernel where every index is already known non-negative that is
wasted work, so the indexing entry points have a u-prefixed twin that skips the
negative-index wrap for runtime signed args — the per-call equivalent of building
with -DTNY_NO_NEGATIVE_INDEX. There are just two:
ugetis the twin ofoperator()— one entry point covering all three forms, chosen by the argument types exactly asoperator()does: all-integer → elementT&, any slice arg → a view, oneellipsis→ expand and re-dispatch.uatis the twin ofat— one element as a rank-0 view.
x.uget(i, j, k); // like x(i,j,k) -> T& (element, no wrap)
x.uget(0, slice(1, 4)); // like x(0, slice(1,4)) -> a VIEW (no wrap on runtime bounds)
x.uget(1, ellipsis, 2); // like x(1, ellipsis, 2) — ellipsis expands, stays unchecked
x.uat(i, j); // like x.at(i,j) -> rank-0 view
They return the same type as their checked counterparts, so they drop in
unchanged, and static (Int<>) bounds and none fold identically — only runtime
signed values skip the wrap. Slice ranges are still clamped, so a runtime negative
bound is now taken literally then clamped (the rule is just "no wrap": with a forward
step a negative stop yields an empty axis, a negative start clamps to 0).
The one rule: passing a negative runtime index to a u-accessor is undefined
behaviour — that is the promise you make in exchange for the tighter codegen.
Bounds checking is off by default (like mdspan's subscript), so normally uget
differs from operator() only by the dropped wrap. Under -DTNY_HARDENED the
checked accessors gain a per-index bounds check; the u* accessors always skip it,
so they are the deliberate opt-out. The bounds check is always off on the device.
Slicing → a sub-view (no copy)¶
If any argument is a slice specifier, operator() returns a lower- or same-rank
view into the same memory:
| argument | meaning |
|---|---|
| an integer | drop that axis (fix it at this index) |
all |
keep the whole axis |
slice(a, b) |
half-open range [a, b) |
slice(a, b, step) |
strided range (step may be negative) |
none (or newaxis) |
inside a slice: an open end. As a bare argument: insert a new size-1 axis |
ellipsis (or etc) |
stand-in for as many all as it takes to fill the rank |
t(1, all, all); // fix axis 0 -> lower-rank view
t(0, slice(1, 4)); // axis 1 range [1,4)
t(all, slice(none, 8, 2)); // every other element up to 8
t(all, slice(none, none, -1)); // reverse an axis (numpy a[:, ::-1])
none is python's None: slice(none, k) starts at 0, slice(m, none) runs to
the end, and slice(none, none) keeps the whole axis (it is all, and folds
to a static extent). Negative bounds wrap.
Ellipsis (..., numpy's Ellipsis) fills the middle for you: it expands to
rank - (#other args) copies of all (possibly zero). At most one per call.
t(1, ellipsis); // == t(1, all, all) drop the FIRST axis, keep the rest
t(ellipsis, 2); // == t(all, all, 2) drop the LAST axis
t(1, ellipsis, 2); // == t(1, all, 2) fill only the middle
t(ellipsis) = b; // whole-tensor slice-assign (copies b's elements in)
t(1, etc, 2); // `etc` is the same marker under another name (== ellipsis)
ellipsis and etc are two names for one marker — "the unspecified middle axes."
etc is the spelling used in an anyshape<etc, …> boundary tag; both
names work in both places.
Inserting a new axis — none / newaxis¶
A bare none argument (not inside a slice) is numpy's newaxis
(a[None] / a[np.newaxis]): it inserts a size-1 axis at that position and
consumes no source axis — the mirror of an integer, which consumes a source
axis and emits nothing. newaxis is just a named alias of none, for
readers who'd rather spell it that way (as far as teeny is concerned they are
the exact same value):
t(none, all, all); // (H,W) -> (1,H,W): insert a leading axis
t(all, all, none); // (H,W) -> (H,W,1): insert a trailing axis
t(newaxis, all, all); // identical to the first line
t(0, none, all); // (H,W) -> (1,W): int drops an axis, none adds one back
t(none, ellipsis); // ellipsis + none compose freely
t(none, ...) at position k is the same view as t.unsqueeze<k>(); the
inserted axis has static extent 1 and stride 0 (it's a broadcast axis, so
writing through it writes the single underlying row/column). Any number of
none/newaxis args can appear alongside integers, ranges, and (at most one)
ellipsis in the same call.
If what remains after expansion is all integers you get an element T&; anything
else yields a view. Assigning into a slice — t(...) = b, t(0, all) = 5.0 —
copies (see Tensors & ownership, which contrasts it with the
rebinding a = b). The copy runs front-to-back with no overlap check, so a slice
that overlaps its own source (a(slice(1,5)) = a(slice(0,4))) reads cells it
has already written — clone() the source first if the regions overlap.
Axes kept with all keep their static extent; a ranged axis becomes dynamic
(its size is a runtime value). Reach for all when you want the extent to keep
folding. The output layout is strides<...> with each kept stride folded to a
compile-time value where derivable.
Runtime vs compile-time bounds. slice(a, b) takes its bounds as ordinary
values; slice<a,b>() / slice<a,b,step>() bake them into the type, so a ranged
axis keeps a static extent and stride (folds like all) instead of a runtime
one. The runtime form is the value spelling; reach for the compile-time form when
you want the extent to keep folding:
take_along<Axes...> — bind named axes¶
operator() is positional. To name only the axes you touch and keep the rest,
use take_along:
Each bind argument is an integer (negatives wrap), all, or a slice — the same
specifiers operator() accepts. The value form leads with an axis<...>{}
selector (a compile-time axis list, sibling of shape<...>, like numpy's
axis: int | list[int]); being a single deduced argument it needs no .template
on a type-dependent receiver, and it's the one spelling that disambiguates
take_along's two argument packs cleanly.