Skip to content

Commit 2adc8b5

Browse files
authored
gh-157710: Move docs for mutating PyUnicode to new section (GH-159028)
1 parent 438c999 commit 2adc8b5

1 file changed

Lines changed: 150 additions & 151 deletions

File tree

‎Doc/c-api/unicode.rst‎

Lines changed: 150 additions & 151 deletions
Original file line numberDiff line numberDiff line change
@@ -154,29 +154,6 @@ access to internal read-only data of Unicode objects:
154154
.. versionadded:: 3.3
155155
156156
157-
.. c:function:: void PyUnicode_WRITE(int kind, void *data, \
158-
Py_ssize_t index, Py_UCS4 value)
159-
160-
Write the code point *value* to the given zero-based *index* in a string.
161-
162-
The *kind* value and *data* pointer must have been obtained from a
163-
string using :c:func:`PyUnicode_KIND` and :c:func:`PyUnicode_DATA`
164-
respectively. You must hold a reference to that string while calling
165-
:c:func:`!PyUnicode_WRITE`. All requirements of
166-
:c:func:`PyUnicode_WriteChar` also apply.
167-
168-
The function performs no checks for any of its requirements,
169-
and is intended for usage in loops.
170-
171-
While :class:`str` objects are usually immutable in Python, this special C API allows
172-
mutating a fresh :class:`str` object if the string has not been "used" yet.
173-
174-
.. versionadded:: 3.3
175-
176-
.. soft-deprecated:: next
177-
Use the :c:type:`PyUnicodeWriter` API instead.
178-
179-
180157
.. c:function:: Py_UCS4 PyUnicode_READ(int kind, void *data, Py_ssize_t index)
181158
182159
Read a code point from a canonical representation *data* (as obtained with
@@ -385,44 +362,6 @@ Creating and accessing Unicode strings
385362
To create Unicode objects and access their basic sequence properties, use these
386363
APIs:
387364
388-
.. c:function:: PyObject* PyUnicode_New(Py_ssize_t size, Py_UCS4 maxchar)
389-
390-
Create a new Unicode object. *maxchar* should be the true maximum code point
391-
to be placed in the string. As an approximation, it can be rounded up to the
392-
nearest value in the sequence 127, 255, 65535, 1114111.
393-
394-
On error, set an exception and return ``NULL``.
395-
396-
After creation, the string can be filled by :c:func:`PyUnicode_WriteChar`,
397-
:c:func:`PyUnicode_CopyCharacters`, :c:func:`PyUnicode_Fill`,
398-
:c:func:`PyUnicode_WRITE` or similar.
399-
Since strings are supposed to be immutable, take care to not “use” the
400-
result while it is being modified. In particular, before it's filled
401-
with its final contents, a string:
402-
403-
- must not be hashed,
404-
- must not be :c:func:`converted to UTF-8 <PyUnicode_AsUTF8AndSize>`,
405-
or another non-"canonical" representation,
406-
- must not have its reference count changed,
407-
- must not be shared with code that might do one of the above.
408-
409-
This list is not exhaustive. Avoiding these uses is your responsibility;
410-
Python does not always check these requirements.
411-
412-
To avoid accidentally exposing a partially-written string object, prefer
413-
using the :c:type:`PyUnicodeWriter` API, or one of the ``PyUnicode_From*``
414-
functions below.
415-
416-
While :class:`str` objects are usually immutable in Python, this special C API
417-
returns a :class:`str` object that can be mutated, except if *size* is zero, in which
418-
case it returns the immutable empty string constant.
419-
420-
.. versionadded:: 3.3
421-
422-
.. soft-deprecated:: next
423-
Use the :c:type:`PyUnicodeWriter` API instead.
424-
425-
426365
.. c:function:: PyObject* PyUnicode_FromKindAndData(int kind, const void *buffer, \
427366
Py_ssize_t size)
428367
@@ -755,96 +694,6 @@ APIs:
755694
.. versionadded:: 3.3
756695
757696
758-
.. c:function:: Py_ssize_t PyUnicode_CopyCharacters(PyObject *to, \
759-
Py_ssize_t to_start, \
760-
PyObject *from, \
761-
Py_ssize_t from_start, \
762-
Py_ssize_t how_many)
763-
764-
Copy characters from one Unicode object into another. This function performs
765-
character conversion when necessary and falls back to :c:func:`!memcpy` if
766-
possible. Returns ``-1`` and sets an exception on error, otherwise returns
767-
the number of copied characters.
768-
769-
While :class:`str` objects are usually immutable in Python, this special C API allows
770-
mutating a fresh :class:`str` object if the string has not been "used" yet.
771-
772-
See :c:func:`PyUnicode_New` for details.
773-
774-
.. versionadded:: 3.3
775-
776-
.. soft-deprecated:: next
777-
Use the :c:type:`PyUnicodeWriter` API instead.
778-
779-
780-
.. c:function:: int PyUnicode_Resize(PyObject **unicode, Py_ssize_t length);
781-
782-
Resize a Unicode object *\*unicode* to the new *length* in code points.
783-
784-
Try to resize the string in place (which is usually faster than allocating
785-
a new string and copying characters), or create a new string.
786-
787-
*\*unicode* is modified to point to the new (resized) object and ``0`` is
788-
returned on success. Otherwise, ``-1`` is returned and an exception is set,
789-
and *\*unicode* is left untouched.
790-
791-
The function doesn't check string content, the result may not be a
792-
string in canonical representation.
793-
794-
While :class:`str` objects are usually immutable in Python, this special C API
795-
can resize a :class:`str` object in-place if the string has not been "used" yet.
796-
It returns a :class:`str` object which can be mutated, except if *size* is zero, in
797-
which case it returns the immutable empty string constant.
798-
799-
.. soft-deprecated:: next
800-
Use the :c:type:`PyUnicodeWriter` API instead.
801-
802-
803-
.. c:function:: Py_ssize_t PyUnicode_Fill(PyObject *unicode, Py_ssize_t start, \
804-
Py_ssize_t length, Py_UCS4 fill_char)
805-
806-
Fill a string with a character: write *fill_char* into
807-
``unicode[start:start+length]``.
808-
809-
Fail if *fill_char* is bigger than the string maximum character, or if the
810-
string has more than 1 reference.
811-
812-
Return the number of written characters, or return ``-1`` and raise an
813-
exception on error.
814-
815-
While :class:`str` objects are usually immutable in Python, this special C API allows
816-
mutating a fresh :class:`str` object if the string has not been "used" yet.
817-
818-
See :c:func:`PyUnicode_New` for details.
819-
820-
.. versionadded:: 3.3
821-
822-
.. soft-deprecated:: next
823-
Use the :c:type:`PyUnicodeWriter` API instead.
824-
825-
826-
.. c:function:: int PyUnicode_WriteChar(PyObject *unicode, Py_ssize_t index, \
827-
Py_UCS4 character)
828-
829-
Write a *character* to the string *unicode* at the zero-based *index*.
830-
Return ``0`` on success, ``-1`` on error with an exception set.
831-
832-
This function checks that *unicode* is a Unicode object, that the index is
833-
not out of bounds, and that the object's reference count is one.
834-
See :c:func:`PyUnicode_WRITE` for a version that skips these checks,
835-
making them your responsibility.
836-
837-
While :class:`str` objects are usually immutable in Python, this special C API allows
838-
mutating a fresh :class:`str` object if the string has not been "used" yet.
839-
840-
See :c:func:`PyUnicode_New` for details.
841-
842-
.. versionadded:: 3.3
843-
844-
.. soft-deprecated:: next
845-
Use the :c:type:`PyUnicodeWriter` API instead.
846-
847-
848697
.. c:function:: Py_UCS4 PyUnicode_ReadChar(PyObject *unicode, Py_ssize_t index)
849698
850699
Read a character from a string. This function checks that *unicode* is a
@@ -2053,3 +1902,153 @@ The following API is deprecated.
20531902
This API does nothing since Python 3.12.
20541903
Previously, this could be called to check if
20551904
:c:func:`PyUnicode_READY` is necessary.
1905+
1906+
1907+
.. _pyunicode-new-mutating:
1908+
1909+
Mutating string objects
1910+
"""""""""""""""""""""""
1911+
1912+
The following functions allow creating a blank string object of a given size,
1913+
then filling in its contents.
1914+
They are :term:`soft deprecated`: they break the assumption that
1915+
strings are immutable, making them hard to use correctly.
1916+
1917+
Prefer using the :c:type:`PyUnicodeWriter` API,
1918+
or one of the ``PyUnicode_From*``
1919+
functions such as :c:func:`PyUnicode_FromStringAndSize`.
1920+
1921+
If you do use the functions below, take care to not "use" such a string while
1922+
it is being modified.
1923+
In particular, before it's filled with its final contents, a string:
1924+
1925+
- must not be hashed,
1926+
- must not be :c:func:`converted to UTF-8 <PyUnicode_AsUTF8AndSize>`,
1927+
or another non-"canonical" representation,
1928+
- must not have its reference count changed,
1929+
- must not be accessed from another thread,
1930+
- must not be shared with code that might do one of the above.
1931+
1932+
This list is not exhaustive. Avoiding these uses is your responsibility;
1933+
Python does not always check these requirements.
1934+
1935+
1936+
.. c:function:: PyObject* PyUnicode_New(Py_ssize_t size, Py_UCS4 maxchar)
1937+
1938+
Create a new Unicode object. *maxchar* should be the true maximum code point
1939+
to be placed in the string. As an approximation, it can be rounded up to the
1940+
nearest value in the sequence 127, 255, 65535, 1114111.
1941+
1942+
On error, set an exception and return ``NULL``.
1943+
1944+
See :ref:`pyunicode-new-mutating` for important warnings and caveats.
1945+
1946+
.. versionadded:: 3.3
1947+
1948+
.. soft-deprecated:: next
1949+
See :ref:`pyunicode-new-mutating`.
1950+
1951+
1952+
.. c:function:: void PyUnicode_WRITE(int kind, void *data, \
1953+
Py_ssize_t index, Py_UCS4 value)
1954+
1955+
Write the code point *value* to the given zero-based *index* in a string.
1956+
1957+
The *kind* value and *data* pointer must have been obtained from a
1958+
string using :c:func:`PyUnicode_KIND` and :c:func:`PyUnicode_DATA`
1959+
respectively. You must hold a reference to that string while calling
1960+
:c:func:`!PyUnicode_WRITE`. All requirements of
1961+
:c:func:`PyUnicode_WriteChar` also apply.
1962+
1963+
The function performs no checks for any of its requirements,
1964+
and is intended for usage in loops.
1965+
1966+
The owning string must not be "used" yet.
1967+
See :ref:`pyunicode-new-mutating` for details.
1968+
1969+
.. versionadded:: 3.3
1970+
1971+
.. soft-deprecated:: next
1972+
Use the :c:type:`PyUnicodeWriter` API instead.
1973+
1974+
1975+
.. c:function:: Py_ssize_t PyUnicode_CopyCharacters(PyObject *to, \
1976+
Py_ssize_t to_start, \
1977+
PyObject *from, \
1978+
Py_ssize_t from_start, \
1979+
Py_ssize_t how_many)
1980+
1981+
Copy characters from one Unicode object into another. This function performs
1982+
character conversion when necessary and falls back to :c:func:`!memcpy` if
1983+
possible. Returns ``-1`` and sets an exception on error, otherwise returns
1984+
the number of copied characters.
1985+
1986+
The destination string must not be "used" yet.
1987+
See :ref:`pyunicode-new-mutating` for details.
1988+
1989+
.. versionadded:: 3.3
1990+
1991+
.. soft-deprecated:: next
1992+
Use the :c:type:`PyUnicodeWriter` API instead.
1993+
1994+
1995+
.. c:function:: int PyUnicode_Resize(PyObject **unicode, Py_ssize_t length);
1996+
1997+
Resize a Unicode object *\*unicode* to the new *length* in code points.
1998+
1999+
Try to resize the string in place (which is usually faster than allocating
2000+
a new string and copying characters), or create a new string.
2001+
2002+
*\*unicode* is modified to point to the new (resized) object and ``0`` is
2003+
returned on success. Otherwise, ``-1`` is returned and an exception is set,
2004+
and *\*unicode* is left untouched.
2005+
2006+
The function doesn't check string content, the result may not be a
2007+
string in canonical representation.
2008+
2009+
*\*unicode* must not be "used" yet.
2010+
See :ref:`pyunicode-new-mutating` for details.
2011+
2012+
.. soft-deprecated:: next
2013+
Use the :c:type:`PyUnicodeWriter` API instead.
2014+
2015+
2016+
.. c:function:: Py_ssize_t PyUnicode_Fill(PyObject *unicode, Py_ssize_t start, \
2017+
Py_ssize_t length, Py_UCS4 fill_char)
2018+
2019+
Fill a string with a character: write *fill_char* into
2020+
``unicode[start:start+length]``.
2021+
2022+
Fail if *fill_char* is bigger than the string maximum character, or if the
2023+
string has more than 1 reference.
2024+
2025+
Return the number of written characters, or return ``-1`` and raise an
2026+
exception on error.
2027+
2028+
*unicode* must not be "used" yet.
2029+
See :ref:`pyunicode-new-mutating` for details.
2030+
2031+
.. versionadded:: 3.3
2032+
2033+
.. soft-deprecated:: next
2034+
Use the :c:type:`PyUnicodeWriter` API instead.
2035+
2036+
2037+
.. c:function:: int PyUnicode_WriteChar(PyObject *unicode, Py_ssize_t index, \
2038+
Py_UCS4 character)
2039+
2040+
Write a *character* to the string *unicode* at the zero-based *index*.
2041+
Return ``0`` on success, ``-1`` on error with an exception set.
2042+
2043+
This function checks that *unicode* is a Unicode object, that the index is
2044+
not out of bounds, and that the object's reference count is one.
2045+
See :c:func:`PyUnicode_WRITE` for a version that skips these checks,
2046+
making them your responsibility.
2047+
2048+
*unicode* must not be "used" yet.
2049+
See :ref:`pyunicode-new-mutating` for details.
2050+
2051+
.. versionadded:: 3.3
2052+
2053+
.. soft-deprecated:: next
2054+
Use the :c:type:`PyUnicodeWriter` API instead.

0 commit comments

Comments
 (0)