why use cell object in python closure implemention?

Viewed 362
def outer():
    n = 1
    def inner():
        return n
    n = 2
    return inner

inner = outer()
print innner()  # output 2

I know it well that how CPython implement the closure, my question is not why output is 2, but why Python design it to output 2.

python use cell object in the closure implementiton, which indirectly ref the exactly PyObject which we want to capture. PythonVM create exactly one cell object for one freevar, in this example, in the outer scope, cell object first ref to 1, then ref to 2. when we call inner function, freevar always load the newest value in the outer function, so output 2.

The "cell object" is the additional abstract level in closure implemention. Actually I modifyed a few lines of the CPython code about the STORE_DEREF and LOAD_DEREF opcode process, remove the "cell object" level, save the real object in the inner's closure. Then the example will output 1. Everything runs ok except a simple traceback in standard library, some code assume cell is hashable. But i think it's not a big matter.

I think output "1" is intuitive sense. So my question is why python make a "cell object" level in the closure implemention ?I know the implemention clearly, but why python design like this ?

1 Answers

The mental model is that the nested function (including a lambda, which is purely a syntactic difference) uses the same variable as the outer function and therefore observes changes to its value even after the function is created. This can be useful: a nested function is always up-to-date with assignments in a long function (that contains and calls it). This also, however, gives the famous issue with lambdas created in a loop: they all share the one loop variable.

This model is no more or less powerful than one where the function captures a value: to emulate that mode, you just create another variable just for the nested function’s use (which means another function call if a loop is involved), while to emulate Python’s behavior with value capture you just capture a container whose (single) element can be mutated.

It’s a matter of philosophy and language consistency as to which behavior is favored by the syntax. The decision here is to have all reads be of a variable; C++, by contrast, supports both behaviors even within one lambda, and even allows (with mutable) updating copies of captured values.

Related