Skip to content

Shape Typing #15

Description

@34j

@jorenham

I feel that perfect Shape typing, especially typing "broadcasting" is impossible (or requires extremely large number of alias), what are your thoughts?

I thought allowing Literal as an argument to the tuple would make typing impossible for some functions, and would better to provide only information on the array dimension (the number of ints in argument of tuple) using Unpack.

Edit: Didn't know this
python/typing#1439

Activity

  1. 34j commented on Jun 4, 2025

    @34j
    Author

    And furthermore, adding shape typing to the first release seems a bit difficult, I suggest making it all Any at first.

  2. jorenham commented on Jun 4, 2025

    @jorenham
    Member

    I feel that perfect Shape typing, especially typing "broadcasting" is impossible (or requires extremely large number of alias), what are your thoughts?

    Ah, shape-typing! Now that's a topic I like to see.

    And coincidentally, I actually figured out a way to express numpy's broadcasting rules in Python's type system in numtype. User's will only ever need to use one of two type aliases: HasRankGT or HasRankLE. But under the hood, I had to do some typing acrobatics and emulate rust-like associated types to get this to work. The big limitation it isn't as general as I'd like it to be, and will only work for the first couple of ndim/ranks. But it's pretty easy to extend number, and it's also easy to use this in downstream libraries. Anyway, it's probably easiest (for you) to look at the implementation and the (exhaustively generated) type-tests if want to know the juicy deets.

    I thought allowing Literal as an argument to the tuple would make typing impossible for some functions, and would better to provide only information on the array dimension (the number of ints in argument of tuple) using Unpack.

    There has been some recent discussion on typing the individual axes of the shapes, and why that's probably not a good idea (even if literal int's would've worked like everyone expects them to): numpy/numtype#578.

    And furthermore, adding shape typing to the first release seems a bit difficult, I suggest making it all Any at first.

    Starting gradual typing indeed sounds like the way to go here.

    But instead of Any, perhaps we could get away with jumping straight to the "gradual tuple" tuple[Any, ...], which numpy recently switched to as default 🤔. See numpy/numpy#28982 an numpy/numtype#573 for the backstory. However, there are some mypy problems that make working with it variadic context a bit, erm, "challenging": https://github.com/python/mypy/issues?q=is%3Aissue%20author%3Ajorenham%20label%3Atopic-pep-646

  3. 34j commented on Jun 5, 2025

    @34j
    Author

    If a type or overload is created for each dimension like this, I suppose the complexity (the number of lines, types or overloads) is probably proportional to the maximum dimension N, but I'm concerned about the performance hit to the IDE or mypy. How large could N be practically? As a reference, Numpy currently seems to support up to 32 dimensions, but I don't see anything specified about this in the array-api?

    Also I was thinking of an overload approach, but due to a specific problem with VSCode displaying all overloads before docstrings, maybe it is better to use your approach for readability?

  4. jorenham commented on Jun 5, 2025

    @jorenham
    Member

    If a type or overload is created for each dimension like this

    For two shaped inputs, you only need two HasRankLE overloads to cover all possible combinations of supported ranks (currently 0-4):

    from typing import Any, TypeVar, overload, reveal_type
    
    import _numtype as _nt
    import numpy as np
    
    _ShapeT = TypeVar("_ShapeT", bound=_nt.Shape)
    
    @overload
    def demo(a: _nt.HasRankLE[_ShapeT], b: np.ndarray[_ShapeT]) -> _nt.Array[Any, _ShapeT]: ...
    @overload
    def demo(a: np.ndarray[_ShapeT], b: _nt.HasRankLE[_ShapeT]) -> _nt.Array[Any, _ShapeT]: ...
    
    a1: np.ndarray[_nt.Rank1]
    reveal_type(demo(a1, a1))  # Type of "demo(a1, a1)" is "ndarray[Rank1, dtype[Any]]"
    
    a2: np.ndarray[_nt.Rank2]
    reveal_type(demo(a2, a2))  # Type of "demo(a2, a2)" is "ndarray[Rank2, dtype[Any]]"
    reveal_type(demo(a1, a2))  # Type of "demo(a1, a2)" is "ndarray[Rank2, dtype[Any]]"
    reveal_type(demo(a2, a1))  # Type of "demo(a2, a1)" is "ndarray[Rank2, dtype[Any]]"
    
    a3: np.ndarray[_nt.Rank3]
    reveal_type(demo(a3, a3))  # Type of "demo(a3, a3)" is "ndarray[Rank3, dtype[Any]]"
    reveal_type(demo(a1, a3))  # Type of "demo(a1, a3)" is "ndarray[Rank3, dtype[Any]]"
    reveal_type(demo(a2, a3))  # Type of "demo(a2, a3)" is "ndarray[Rank3, dtype[Any]]"
    reveal_type(demo(a3, a1))  # Type of "demo(a3, a1)" is "ndarray[Rank3, dtype[Any]]"
    reveal_type(demo(a3, a2))  # Type of "demo(a3, a2)" is "ndarray[Rank3, dtype[Any]]"
    
    # Rank4 also works
  5. jorenham commented on Jun 5, 2025

    @jorenham
    Member

    Numpy currently seems to support up to 32 dimensions

    Since numpy 2 this has been increased to 64 😅. But yea, my assumption here is that in 99% of use-cases won't go higher than rank 4. And I based that on totally nothing 💩.

  6. 34j commented on Jun 5, 2025

    @34j
    Author

    Hmmm wonderful work but it is too sophisticated for me to understand. So since this is not allowed

    @overload
    def demo(a: NDArray[tuple[*T2]], b: NDArray[tuple[*T1, *T2]]) -> NDArray[tuple[*T1, *T2]]:
        ....
    
    @overload
    def demo(a: NDArray[tuple[*T1, *T2]], b: NDArray[tuple[*T2]]) -> NDArray[tuple[*T1, *T2]]:
        ....

    you have introduced functions __broadcast__ and __own_shape__ into tuple that don't actually exist in order to represent the relationship between tuple which is now Rank? I am most confused about the role of __own_shape__. Is there an explanation somewhere? What is _OwnShapeT_contra and _OwnShapeT_co?

  7. 34j commented on Jun 5, 2025

    @34j
    Author

    Isn't this just defining magnitude relationship between the number of arguments in a tuple[int, ...]? There is no need to limit the argument to int, and this is no longer related to the Array API, therefore this could also be a single independent package, isn't it?

  8. jorenham commented on Jun 5, 2025

    @jorenham
    Member

    you have introduced functions __broadcast__ and __own_shape__ into tuple that don't actually exist in order to represent the relationship between tuple which is now Rank? I am most confused about the role of __own_shape__. Is there an explanation somewhere? What is _OwnShapeT_contra and _OwnShapeT_co?

    No explanation yet I'm afraid. This is all pretty experimental at the moment. But I plan to document it at https://numpy.org/numtype/ at some point.

    There's a lot of trickey going on here, and I often have trouble wrapping my head around how it actually works. But the intuition behind those type-check-only methods, is that they encode the broadcasting rules, i.e. which ranks it can broadcast to, and the ranks that can broadcast to itself (the from_ param).
    They're in a contravariant position so that the HasRank aliases work with both tuple and Rank* type-args. The "matching" is all flipped upside down, which makes it very difficult to reason about.

    And the __own_shape__ is basically used as workaround for the lack of a Super[T] type that returns the supertype of T, which in this case is the original tuple type.

  9. jorenham commented on Jun 5, 2025

    @jorenham
    Member

    Isn't this just defining magnitude relationship between the number of arguments in a tuple[int, ...]?

    Yea exactly. This is basically an extremely overengineered implementation of min and max.

  10. jorenham commented on Jun 5, 2025

    @jorenham
    Member

    There is no need to limit the argument to int, and this is no longer related to the Array API, therefore this could also be a single independent package, isn't it?

    This will only work if all of numpy's (constructor) functions will return arrays with a Rank* shape-type, instead of a plain tuple, because it's not possible to statically convert a tuple to a Rank I'm afraid.

    And theoretically it's possible to use subtypes of ints. But I don't think that there exists a problem that can be solved by doing so.

  11. jorenham commented on Jun 5, 2025

    @jorenham
    Member

    @RBerga06, in case you want to know more about the numtype shape-typing implementation, there's a chance that my ramblings above could help with that :)

  12. 34j commented on Jun 5, 2025

    @34j
    Author

    Hmmm, so given that we successfully np.ndarray[TShape, TDtype] = ArrayAPINamespace[np.Array, TShape, TDtype, Device(?)], it would be like:

    • If jorenham's method is used, if a user has an NDArray[tuple[int]], it needs to be converted to NDArray[Rank0] to have Shape typing
    • If we did similar to
    @overload
    def demo(a: NDArray[tuple[*TE]], b: NDArray[tuple[*TE]]) -> NDArray[tuple[*TE]]:
        ....
    
    @overload
    def demo(a: NDArray[tuple[*TE]], b: NDArray[tuple[int, *TE]]) -> NDArray[tuple[int, *TE]]:
        ....
    
    
    @overload
    def demo(a: NDArray[tuple[*TE]], b: NDArray[tuple[int, int, *TE]]) -> NDArray[tuple[int, int, *TE]]:
        ....
    
    
    @overload
    def demo(a: NDArray[tuple[*TE]], b: NDArray[tuple[int, int, int, *TE]]) -> NDArray[tuple[int, int, int, *TE]]:
        ....
    

    user does not need to convert it?

  13. jorenham commented on Jun 5, 2025

    @jorenham
    Member
    • If jorenham's method is used, if a user has an NDArray[tuple[int]], it needs to be converted to NDArray[Rank0] to have Shape typing

    FWIW; only one of the arrays involved needs to have a Rank* shape-type with my method. And to help with that, I'm also planning providing (runtime-accessible) shaped array-type aliases to help with that, such as numtype.Array2, or equivalently, numtype.Array[T, numtype.Rank2]. So the numtype package and the new numpy-stubs stubs will both be installed when you uv pip install numtype (in the near future).


    And your overload approach is indeed what I've currently been doing in numpy and scipy-stubs. But in many cases that'll lead to too many overloads (personally, I draw the line at >64).

  14. jorenham commented on Jun 8, 2025

    @jorenham
    Member

    shape-typing support is not completed yet

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions