Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
45 changes: 40 additions & 5 deletions Doc/c-api/unicode.rst
Original file line number Diff line number Diff line change
Expand Up @@ -168,8 +168,14 @@ access to internal read-only data of Unicode objects:
The function performs no checks for any of its requirements,
and is intended for usage in loops.

While str objects are usually immutable in Python, this special C API allows
mutating a fresh str object if the string was not “used” yet.

.. versionadded:: 3.3

.. soft-deprecated:: next
Use the :c:type:`PyUnicodeWriter` API instead.


.. c:function:: Py_UCS4 PyUnicode_READ(int kind, void *data, Py_ssize_t index)

Expand Down Expand Up @@ -407,9 +413,15 @@ APIs:
using the :c:type:`PyUnicodeWriter` API, or one of the ``PyUnicode_From*``
functions below.

While str objects are usually immutable in Python, this special C API
returns a str object which can be mutated; except if *size* is zero in which
case it returns the immutable empty string.

.. versionadded:: 3.3

.. soft-deprecated:: next
Use the :c:type:`PyUnicodeWriter` API instead.


.. c:function:: PyObject* PyUnicode_FromKindAndData(int kind, const void *buffer, \
Py_ssize_t size)
Expand Down Expand Up @@ -754,11 +766,16 @@ APIs:
possible. Returns ``-1`` and sets an exception on error, otherwise returns
the number of copied characters.

The string must not have been “used” yet.
While str objects are usually immutable in Python, this special C API allows
mutating a fresh str object if the string was not “used” yet.

See :c:func:`PyUnicode_New` for details.

.. versionadded:: 3.3

.. soft-deprecated:: next
Use the :c:type:`PyUnicodeWriter` API instead.


.. c:function:: int PyUnicode_Resize(PyObject **unicode, Py_ssize_t length);

Expand All @@ -774,6 +791,14 @@ APIs:
The function doesn't check string content, the result may not be a
string in canonical representation.

While str objects are usually immutable in Python, this special C API
can resize a str object in-place if the string was not “used” yet.
It returns a str object which can be mutated; except if *size* is zero in
which case it returns the immutable empty string.

.. soft-deprecated:: next
Use the :c:type:`PyUnicodeWriter` API instead.


.. c:function:: Py_ssize_t PyUnicode_Fill(PyObject *unicode, Py_ssize_t start, \
Py_ssize_t length, Py_UCS4 fill_char)
Expand All @@ -784,14 +809,19 @@ APIs:
Fail if *fill_char* is bigger than the string maximum character, or if the
string has more than 1 reference.

The string must not have been “used” yet.
See :c:func:`PyUnicode_New` for details.

Return the number of written characters, or return ``-1`` and raise an
exception on error.

While str objects are usually immutable in Python, this special C API allows
mutating a fresh str object if the string was not “used” yet.

See :c:func:`PyUnicode_New` for details.

.. versionadded:: 3.3

.. soft-deprecated:: next
Use the :c:type:`PyUnicodeWriter` API instead.


.. c:function:: int PyUnicode_WriteChar(PyObject *unicode, Py_ssize_t index, \
Py_UCS4 character)
Expand All @@ -804,11 +834,16 @@ APIs:
See :c:func:`PyUnicode_WRITE` for a version that skips these checks,
making them your responsibility.

The string must not have been “used” yet.
While str objects are usually immutable in Python, this special C API allows
mutating a fresh str object if the string was not “used” yet.

See :c:func:`PyUnicode_New` for details.

.. versionadded:: 3.3

.. soft-deprecated:: next
Use the :c:type:`PyUnicodeWriter` API instead.


.. c:function:: Py_UCS4 PyUnicode_ReadChar(PyObject *unicode, Py_ssize_t index)

Expand Down
7 changes: 7 additions & 0 deletions Doc/whatsnew/3.16.rst
Original file line number Diff line number Diff line change
Expand Up @@ -1224,6 +1224,13 @@ Deprecated C APIs
:c:func:`PyModule_GetFilenameObject` instead is still recommended.
(Contributed by Victor Stinner in :gh:`154757`.)

* Soft deprecate functions modifying Unicode strings:
:c:func:`PyUnicode_New`, :c:func:`PyUnicode_CopyCharacters`,
:c:func:`PyUnicode_Fill`, :c:func:`PyUnicode_Resize`,
:c:func:`PyUnicode_WRITE` and :c:func:`PyUnicode_WriteChar`.
Use the safer :c:type:`PyUnicodeWriter` API instead.
(Contributed by Victor Stinner in :gh:`157710`.)

.. Add C API deprecations above alphabetically, not here at the end.

.. include:: ../deprecations/c-api-pending-removal-in-3.18.rst
Expand Down
58 changes: 50 additions & 8 deletions Lib/test/test_capi/test_unicode.py
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,7 @@
# Maximum invalid character which fits into 32-bit Py_UCS4
MAX_INVALID_CHAR = 0xFFFF_FFFF
NULL = None
USED_STR_ERROR = 'Cannot modify a string currently used'

class Str(str):
pass
Expand Down Expand Up @@ -76,9 +77,27 @@ def test_checkexact(self):
# Test PyUnicode_CheckExact()
self._test_check(_testlimitedcapi.unicode_checkexact, exact=True)

def assert_is_mutable(self, result):
# Check that result is a "mutable" Unicode string
self.assertEqual(sys.getrefcount(result), 1)
self.assertFalse(sys._is_immortal(result))

def assert_is_empty_singleton(self, result):
# Check that result is the empty string singleton
self.assertEqual(result, '')
self.assertTrue(sys._is_immortal(result))

def test_new(self):
"""Test PyUnicode_New()"""
new = _testcapi.unicode_new
_unicode_new = _testcapi.unicode_new

def new(size, maxchar):
result = _unicode_new(size, maxchar)
if size != 0:
self.assert_is_mutable(result)
else:
self.assert_is_empty_singleton(result)
return result

for maxchar in 0, 0x61, 0xa1, 0x4f60, 0x1f600, 0x10ffff:
self.assertEqual(new(0, maxchar), '')
Expand Down Expand Up @@ -123,6 +142,10 @@ def test_fill(self):
self.assertEqual(fill(to, start, length, fill_char),
(expected, filled))

# A string with 2 references cannot be modified
with self.assertRaisesRegex(SystemError, USED_STR_ERROR):
fill('abc', 0, 3, ord('x'), incref=True)

s = strings[0]
self.assertRaises(IndexError, fill, s, -1, 0, 0x78)
self.assertRaises(IndexError, fill, s, PY_SSIZE_T_MIN, 0, 0x78)
Expand Down Expand Up @@ -162,27 +185,41 @@ def _test_writechar(self, writechar, *, check):

def test_writechar(self):
"""Test PyUnicode_WriteChar()"""
self._test_writechar(_testlimitedcapi.unicode_writechar, check=True)
writechar = _testlimitedcapi.unicode_writechar
self._test_writechar(writechar, check=True)

# A string with 2 references cannot be modified
with self.assertRaisesRegex(SystemError, USED_STR_ERROR):
writechar('abc', 1, ord('x'), incref=True)

def test_write_macro(self):
"""Test PyUnicode_WRITE()"""
self._test_writechar(_testcapi.unicode_write, check=False)

def test_resize(self):
"""Test PyUnicode_Resize()"""
resize = _testlimitedcapi.unicode_resize
_unicode_resize = _testlimitedcapi.unicode_resize

def resize(s, size):
result, int_result = _unicode_resize (s, size)
self.assertEqual(int_result, 0)
if size != 0:
self.assert_is_mutable(result)
else:
self.assert_is_empty_singleton(result)
return result

strings = [
# all strings have exactly 3 characters
'abc', '\xa1\xa2\xa3', '\u4f60\u597d\u4e16',
'\U0001f600\U0001f601\U0001f602'
]
for s in strings:
self.assertEqual(resize(s, 3), (s, 0))
self.assertEqual(resize(s, 2), (s[:2], 0))
self.assertEqual(resize(s, 4), (s + '\0', 0))
self.assertEqual(resize(s, 10), (s + '\0'*7, 0))
self.assertEqual(resize(s, 0), ('', 0))
self.assertEqual(resize(s, 3), s)
self.assertEqual(resize(s, 2), s[:2])
self.assertEqual(resize(s, 4), s + '\0')
self.assertEqual(resize(s, 10), s + '\0'*7)
self.assertEqual(resize(s, 0), '')
self.assertRaises(MemoryError, resize, s, PY_SSIZE_T_MAX)
self.assertRaises(SystemError, resize, s, -1)
self.assertRaises(SystemError, resize, s, PY_SSIZE_T_MIN)
Expand Down Expand Up @@ -1746,6 +1783,11 @@ def test_copycharacters(self):
self.assertRaises(SystemError, unicode_copycharacters, s, 0, s, 0, PY_SSIZE_T_MIN)
self.assertRaises(SystemError, unicode_copycharacters, s, 0, b'', 0, 0)
self.assertRaises(SystemError, unicode_copycharacters, s, 0, [], 0, 0)

# A string with 2 references cannot be modified
with self.assertRaisesRegex(SystemError, USED_STR_ERROR):
unicode_copycharacters('abc', 0, 'abc', 0, 1, incref=True)

# CRASHES unicode_copycharacters(s, 0, NULL, 0, 0)
# TODO: Test PyUnicode_CopyCharacters() with non-unicode and
# non-modifiable unicode as "to".
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
Soft deprecate functions modifying Unicode strings:
:c:func:`PyUnicode_New`, :c:func:`PyUnicode_CopyCharacters`,
:c:func:`PyUnicode_Fill`, :c:func:`PyUnicode_Resize`,
:c:func:`PyUnicode_WRITE` and :c:func:`PyUnicode_WriteChar`. Use
the :c:type:`PyUnicodeWriter` API instead. Patch by Victor Stinner.
40 changes: 33 additions & 7 deletions Modules/_testcapi/unicode.c
Original file line number Diff line number Diff line change
Expand Up @@ -59,13 +59,20 @@ unicode_copy(PyObject *unicode)

/* Test PyUnicode_Fill() */
static PyObject *
unicode_fill(PyObject *self, PyObject *args)
unicode_fill(PyObject *self, PyObject *args, PyObject *kwargs)
{
static char *kwlist[] = {"to", "start", "length", "fill_char",
"incref", NULL};
PyObject *to, *to_copy;
Py_ssize_t start, length, filled;
unsigned int fill_char;
int incref = 0;

if (!PyArg_ParseTuple(args, "OnnI", &to, &start, &length, &fill_char)) {
if (!PyArg_ParseTupleAndKeywords(args, kwargs,
"OnnI|p", kwlist,
&to, &start, &length,
&fill_char, &incref))
{
return NULL;
}

Expand All @@ -74,7 +81,14 @@ unicode_fill(PyObject *self, PyObject *args)
return NULL;
}

if (incref) {
Py_INCREF(to_copy);
}
filled = PyUnicode_Fill(to_copy, start, length, (Py_UCS4)fill_char);
if (incref) {
Py_DECREF(to_copy);
}

if (filled == -1 && PyErr_Occurred()) {
Py_DECREF(to_copy);
return NULL;
Expand Down Expand Up @@ -190,13 +204,18 @@ unicode_asutf8(PyObject *self, PyObject *args)

/* Test PyUnicode_CopyCharacters() */
static PyObject *
unicode_copycharacters(PyObject *self, PyObject *args)
unicode_copycharacters(PyObject *self, PyObject *args, PyObject *kwargs)
{
static char *kwlist[] = {"to", "to_start", "from", "from_start",
"howmany", "incref", NULL};
PyObject *from, *to, *to_copy;
Py_ssize_t from_start, to_start, how_many, copied;
int incref = 0;

if (!PyArg_ParseTuple(args, "UnOnn", &to, &to_start,
&from, &from_start, &how_many)) {
if (!PyArg_ParseTupleAndKeywords(args, kwargs,
"UnOnn|p", kwlist,
&to, &to_start, &from, &from_start,
&how_many, &incref)) {
return NULL;
}

Expand All @@ -210,8 +229,15 @@ unicode_copycharacters(PyObject *self, PyObject *args)
return NULL;
}

if (incref) {
Py_INCREF(to_copy);
}
copied = PyUnicode_CopyCharacters(to_copy, to_start, from,
from_start, how_many);
if (incref) {
Py_DECREF(to_copy);
}

if (copied == -1 && PyErr_Occurred()) {
Py_DECREF(to_copy);
return NULL;
Expand Down Expand Up @@ -856,12 +882,12 @@ static PyType_Spec Writer_spec = {

static PyMethodDef TestMethods[] = {
{"unicode_new", unicode_new, METH_VARARGS},
{"unicode_fill", unicode_fill, METH_VARARGS},
{"unicode_fill", _PyCFunction_CAST(unicode_fill), METH_VARARGS | METH_KEYWORDS},
{"unicode_fromkindanddata", unicode_fromkindanddata, METH_VARARGS},
{"unicode_asucs4", unicode_asucs4, METH_VARARGS},
{"unicode_asucs4copy", unicode_asucs4copy, METH_VARARGS},
{"unicode_asutf8", unicode_asutf8, METH_VARARGS},
{"unicode_copycharacters", unicode_copycharacters, METH_VARARGS},
{"unicode_copycharacters", _PyCFunction_CAST(unicode_copycharacters), METH_VARARGS | METH_KEYWORDS},
{"unicode_GET_CACHED_HASH", unicode_GET_CACHED_HASH, METH_O},
{"test_py_identifier", test_py_identifier, METH_NOARGS},
{"corrupt_unicode", corrupt_unicode, METH_VARARGS},
Expand Down
18 changes: 15 additions & 3 deletions Modules/_testlimitedcapi/unicode.c
Original file line number Diff line number Diff line change
Expand Up @@ -140,14 +140,19 @@ unicode_copy(PyObject *unicode)

/* Test PyUnicode_WriteChar() */
static PyObject *
unicode_writechar(PyObject *self, PyObject *args)
unicode_writechar(PyObject *self, PyObject *args, PyObject *kwargs)
{
static char *kwlist[] = {"to", "index", "character", "incref", NULL};
PyObject *to, *to_copy;
Py_ssize_t index;
unsigned int character;
int result;
int incref = 0;

if (!PyArg_ParseTuple(args, "OnI", &to, &index, &character)) {
if (!PyArg_ParseTupleAndKeywords(args, kwargs,
"OnI|p", kwlist,
&to, &index, &character, &incref))
{
return NULL;
}

Expand All @@ -156,7 +161,14 @@ unicode_writechar(PyObject *self, PyObject *args)
return NULL;
}

if (incref) {
Py_INCREF(to_copy);
}
result = PyUnicode_WriteChar(to_copy, index, (Py_UCS4)character);
if (incref) {
Py_DECREF(to_copy);
}

if (result == -1 && PyErr_Occurred()) {
Py_DECREF(to_copy);
return NULL;
Expand Down Expand Up @@ -1880,7 +1892,7 @@ static PyMethodDef TestMethods[] = {
test_unicode_compare_with_ascii, METH_NOARGS},
{"test_string_from_format", test_string_from_format, METH_NOARGS},
{"test_widechar", test_widechar, METH_NOARGS},
{"unicode_writechar", unicode_writechar, METH_VARARGS},
{"unicode_writechar", _PyCFunction_CAST(unicode_writechar), METH_VARARGS | METH_KEYWORDS},
{"unicode_resize", unicode_resize, METH_VARARGS},
{"unicode_append", unicode_append, METH_VARARGS},
{"unicode_appendanddel", unicode_appendanddel, METH_VARARGS},
Expand Down
Loading