Feature or enhancement
Proposal:
Versions
CPython 3.15.0b4, Ubuntu 24.04.4 LTS, gcc 13.3.0
Enhancement
I noticed that the specialization function for TO_BOOL and the guard used by TO_BOOL_INT accept different sets of integers.
_Py_Specialize_ToBool() selects TO_BOOL_INT for any exact integer:
if (PyLong_CheckExact(value)) {
specialized_op = TO_BOOL_INT;
goto success;
}
However, the TO_BOOL_INT macro uses _GUARD_TOS_INT, and that requires the integer to be compact:
op(_GUARD_TOS_INT, (value -- value)) {
PyObject *value_o = PyStackRef_AsPyObjectBorrow(value);
EXIT_IF(!_PyLong_CheckExactAndCompact(value_o));
}
As a result, a non-compact exact integer can cause TO_BOOL_INT to be selected and then immediately miss its guard.
This mismatch appears to have been introduced by GH-143759. Before the refactoring, TO_BOOL_INT had its own PyLong_CheckExact() guard. The refactoring replaced it with a macro using the shared _GUARD_TOS_INT, whose domain is narrower.
I see two alternative ways to fix this.
Option 1: narrow the specialization function
One option would be to change the integer check in _Py_Specialize_ToBool() so that it agrees with the existing guard:
if (_PyLong_CheckExactAndCompact(value)) {
specialized_op = TO_BOOL_INT;
goto success;
}
With this change, non-compact integers remain on the generic TO_BOOL path instead of repeatedly entering and missing TO_BOOL_INT.
My main concern with this option was whether the additional compactness check in the specialization function could regress the common case (compact int), so I benchmarked both compact and non-compact integers.
The benchmark target contained 100 TO_BOOL sites and was warmed up before each measurement:
start = time.perf_counter_ns()
for _ in range(10_000):
target(value)
elapsed = time.perf_counter_ns() - start
Positive values mean that the patched build was faster:
| Version |
Input |
Performance change |
| CPython 3.15 |
non-compact exact int |
+2.21% (95% CI: +1.64% to +2.86%) |
| CPython 3.15 |
compact exact int |
+0.31% (95% CI: −0.15% to +0.70%) |
| CPython main |
non-compact exact int |
+4.03% (95% CI: +1.43% to +7.10%) |
| CPython main |
compact exact int |
−0.23% (95% CI: −0.50% to +0.07%) |
The non-compact case improved because it no longer repeatedly enters and misses TO_BOOL_INT. For compact integers, both confidence intervals include zero, so I did not find evidence that the stronger specialization check causes a regression.
Option 2: give TO_BOOL_INT an exact-int guard
The other option is to keep _Py_Specialize_ToBool() unchanged and add a guard that checks exact type without requiring compactness:
op(_GUARD_TOS_EXACT_INT, (value -- value)) {
PyObject *value_o = PyStackRef_AsPyObjectBorrow(value);
EXIT_IF(!PyLong_CheckExact(value_o));
}
The TO_BOOL_INT macro would then use the new guard:
macro(TO_BOOL_INT) =
_GUARD_TOS_EXACT_INT +
unused/1 +
unused/2 +
_TO_BOOL_INT +
_POP_TOP_INT;
This would restore the specialization domain from before GH-143759.
I am not sure which of these two options is preferable, but the current mismatch seems worth fixing, so I am opening this issue to get feedback on which direction would be better.
Has this already been discussed elsewhere?
No response given
Links to previous discussion of this feature:
No response
Linked PRs
Feature or enhancement
Proposal:
Versions
CPython 3.15.0b4, Ubuntu 24.04.4 LTS, gcc 13.3.0
Enhancement
I noticed that the specialization function for
TO_BOOLand the guard used byTO_BOOL_INTaccept different sets of integers._Py_Specialize_ToBool()selectsTO_BOOL_INTfor any exact integer:However, the
TO_BOOL_INTmacro uses_GUARD_TOS_INT, and that requires the integer to be compact:As a result, a non-compact exact integer can cause
TO_BOOL_INTto be selected and then immediately miss its guard.This mismatch appears to have been introduced by GH-143759. Before the refactoring,
TO_BOOL_INThad its ownPyLong_CheckExact()guard. The refactoring replaced it with a macro using the shared_GUARD_TOS_INT, whose domain is narrower.I see two alternative ways to fix this.
Option 1: narrow the specialization function
One option would be to change the integer check in
_Py_Specialize_ToBool()so that it agrees with the existing guard:With this change, non-compact integers remain on the generic
TO_BOOLpath instead of repeatedly entering and missingTO_BOOL_INT.My main concern with this option was whether the additional compactness check in the specialization function could regress the common case (compact int), so I benchmarked both compact and non-compact integers.
The benchmark target contained 100
TO_BOOLsites and was warmed up before each measurement:Positive values mean that the patched build was faster:
The non-compact case improved because it no longer repeatedly enters and misses
TO_BOOL_INT. For compact integers, both confidence intervals include zero, so I did not find evidence that the stronger specialization check causes a regression.Option 2: give
TO_BOOL_INTan exact-int guardThe other option is to keep
_Py_Specialize_ToBool()unchanged and add a guard that checks exact type without requiring compactness:The
TO_BOOL_INTmacro would then use the new guard:This would restore the specialization domain from before GH-143759.
I am not sure which of these two options is preferable, but the current mismatch seems worth fixing, so I am opening this issue to get feedback on which direction would be better.
Has this already been discussed elsewhere?
No response given
Links to previous discussion of this feature:
No response
Linked PRs
TO_BOOL_INT#155531