Python Strings
Strings are the most commonly used data type in Python. We can create strings using quotation marks ('or") to create strings.
Creating a string is simple: just assign a value to a variable. For example:
var1 = 'Hello World!' var2 = "Python Example"
Python Accessing Values in Strings
Python does not support a single-character type; a single character is also used as a string in Python.
To access substrings in Python, you can use square brackets to slice strings, as in the following example:
Example (Python 2.0+)
The execution result of the above example is:
var1[0]: H var2[1:5]: ytho
Python String Concatenation
We can slice strings and concatenate them with other strings, as in the following example:
Example (Python 2.0+)
The execution result of the above example
输出 :- Hello Example!
Python Escape Characters
When you need to use special characters in a string, Python uses a backslash (\escape characters. The table below:
| Escape Character | Description |
|---|---|
| \(at end of line) | Line continuation |
| \\ | Backslash symbol |
| \' | Single quote |
| \" | Double quote |
| \a | Bell |
| \b | Backspace |
| \e | Escape |
| \000 | Empty |
| \n | Newline |
| \v | Vertical tab |
| \t | Horizontal tab |
| \r | Carriage return |
| \f | Form feed |
| \oyy | Octal number, y represents a character from 0~7, for example: \012 represents a newline. |
| \xyy | Hexadecimal number, starts with \x, yy represents the character, for example: \x0a represents a newline |
| \other | Other characters are output in normal format |
Python String Operators
In the table below, the instance variable a has the value string "Hello", and variable b has the value "Python":
| Operator | Description | Instance |
|---|---|---|
| + | String concatenation |
>>>a + b
'HelloPython'
|
| * | Repeat output string |
>>>a * 2
'HelloHello'
|
| [] | Get character in string by index |
>>>a[1]
'e'
|
| [ : ] | Slice a portion of the string |
>>>a[1:4]
'ell'
|
| in | Membership operator - Returns True if the string contains the given character |
>>>"H" in a
True
|
| not in | Membership operator - Returns True if the string does not contain the given character |
>>>"M" not in a
True
|
| r/R | Raw string - Raw string: all strings are used directly according to their literal meaning, without escaping special or unprintable characters. Except for adding the letter "r" (uppercase or lowercase) before the first quote of the string, raw strings have almost exactly the same syntax as ordinary strings. | >>>print r'\n'
\n
>>> print R'\n'
\n |
| % | Format string | See the next section |
Example (Python 2.0+)
The execution result of the above program is:
a + b 输出结果: HelloPython a * 2 输出结果: HelloHello a[1] 输出结果: e a[1:4] 输出结果: ell H 在变量 a 中 M 不在变量 a 中 \n \n
Python String Formatting
Python supports formatted string output. Although this may involve very complex expressions, the most basic usage is to insert a value into a string with the string format character %s.
In Python, string formatting uses the same syntax as the sprintf function in C.
The following is an example:
#!/usr/bin/python
print "My name is %s and weight is %d kg!" % ('Zara', 21)
The output result of the above example is:
My name is Zara and weight is 21 kg!
Python string formatting symbols:
| Symbol | Description |
|---|---|
| %c | Format a character and its ASCII code |
| %s | Format a string |
| %d | Format an integer |
| %u | Format an unsigned integer |
| %o | Format an unsigned octal number |
| %x | Format an unsigned hexadecimal number |
| %X | Format an unsigned hexadecimal number (uppercase) |
| %f | Format a floating-point number, can specify precision after the decimal point |
| %e | Format a floating-point number using scientific notation |
| %E | Same as %e, formats a floating-point number using scientific notation |
| %g | Shorthand for %f and %e |
| %G | Shorthand for %F and %E |
| %p | Format the address of a variable using a hexadecimal number |
Formatting operator auxiliary directives:
| Symbol | Function |
|---|---|
| * | Define width or decimal point precision |
| - | Used for left alignment |
| + | Display a plus sign (+) before positive numbers |
| <sp> | Display a space before positive numbers |
| # | Display '0' before octal numbers, and '0x' or '0X' before hexadecimal numbers (depending on whether 'x' or 'X' is used) |
| 0 | Pad with '0' before displayed numbers instead of the default spaces |
| % | '%%' outputs a single '%' |
| (var) | Mapping variables (dictionary arguments) |
| m.n. | m is the minimum total width displayed, n is the number of digits after the decimal point (if available) |
Starting from Python 2.6, a new string formatting function was addedstr.format(), which enhances the string formatting capability.
Python Triple Quotes
In Python, triple quotes can be used to assign complex strings.
Python triple quotes allow a string to span multiple lines. The string can contain newlines, tabs, and other special characters.
The syntax of triple quotes is a pair of consecutive single quotes or double quotes (usually used in pairs).
>>> hi = '''hi there''' >>> hi # repr() 'hi\nthere' >>> print hi # str() hi there
Triple quotes free programmers from the mire of quotes and special strings, keeping a small block of string in the so-called WYSIWYG (What You See Is What You Get) format from beginning to end.
A typical use case is when you need a block of HTML or SQL; at this time you should use triple quotes to mark it, as using the traditional escape character system would be very tedious.
errHTML = '''
<HTML><HEAD><TITLE>
Friends CGI Demo</TITLE></HEAD>
<BODY><H3>ERROR</H3>
<B>%s</B><P>
<FORM><INPUT TYPE=button VALUE=Back
ONCLICK="window.history.back()"></FORM>
</BODY></HTML>
'''
cursor.execute('''
CREATE TABLE users (
login VARCHAR(8),
uid INTEGER,
prid INTEGER)
''')
Unicode Strings
In Python, defining a Unicode string is as simple as defining an ordinary string:
>>> u'Hello World !' u'Hello World !'
The lowercase "u" before the quotation marks indicates that a Unicode string is created here. If you want to add a special character, you can use Python's Unicode-Escape encoding. As shown in the following example:
>>> u'Hello\u0020World !' u'Hello World !'
The replaced \u0020 notation indicates inserting the Unicode character with encoding value 0x0020 (space) at the given position.
Python Built-in String Functions
String methods were gradually added from Python 1.6 to 2.0 — they were also added to Jython.
These methods implement most of the methods in the string module. The table below lists the currently supported built-in string methods. All methods include Unicode support, and some are even specifically for Unicode.
| Method | Description |
|---|---|
|
Capitalize the first character of the string |
|
|
Return a new string with the original string centered and padded with spaces to length width |
|
|
Return the number of occurrences of str in string; if beg or end is specified, return the number of occurrences of str in the specified range |
|
|
Decode string using the encoding specified by encoding; if an error occurs, a ValueError exception is reported by default, unless errors is specified as 'ignore' or 'replace' |
|
|
Encode string using the encoding specified by encoding; if an error occurs, a ValueError exception is reported by default, unless errors is specified as 'ignore' or 'replace' |
|
|
Check whether the string ends with obj. If beg or end is specified, check whether it ends with obj within the specified range. If so, return True; otherwise return False. |
|
|
Convert tab symbols in the string to spaces; the default number of spaces for a tab symbol is 8. |
|
|
Detect whether str is contained in string; if beg and end specify a range, check whether it is contained in the specified range; if so, return the starting index value, otherwise return -1. |
|
|
Format string |
|
|
Same as the find() method, except that it raises an exception if str is not in string. |
|
|
If string has at least one character and all characters are letters or digits, then return True, otherwise return False. |
|
|
If string has at least one character and all characters are letters, return True, otherwise return False. |
|
|
If string contains only decimal digits, return True; otherwise return False. |
|
|
If string contains only digits, return True; otherwise return False. |
|
|
If string contains at least one case-sensitive character, and all such case-sensitive characters are lowercase, return True; otherwise return False. |
|
|
If string contains only numeric characters, return True; otherwise return False. |
|
|
If string contains only whitespace, return True; otherwise return False. |
|
|
If string is titlecased (see title()), return True; otherwise return False. |
|
|
If string contains at least one case-sensitive character, and all such case-sensitive characters are uppercase, return True; otherwise return False. |
|
|
With string as the separator, join all elements (their string representations) in seq into a new string. |
|
|
Return a new string with the original string left-aligned and padded with spaces to length width. |
|
|
Convert all uppercase characters in string to lowercase. |
|
|
Strip the leading spaces from string. |
|
|
The maketrans() method is used to create a translation table for character mapping. For the simplest calling method that accepts two parameters, the first parameter is a string representing the characters to be converted, and the second parameter is also a string representing the conversion target. |
|
|
Return thestrlargest letter in the string. |
|
|
Return thestrsmallest letter in the string. |
|
|
Somewhat like a combination of find() and split(); starting from the first position where str appears, split the string into a 3-element tuple (string_pre_str, str, string_post_str). If str is not contained in string, then string_pre_str == string. |
|
|
Replace str1 in string with str2; if num is specified, replace no more than num times. |
|
|
Similar to the find() function; returns the position of the last occurrence of the string, or -1 if there is no match. |
|
|
Similar to index(), but returns the index of the last matched substring. |
|
|
Return a new string with the original string right-aligned and padded with spaces to length width. |
|
|
Similar to the partition() function, but searches from the right. |
|
|
Strip the trailing spaces from string. |
|
|
With str as the delimiter, slice string; if num has a specified value, only splitnum+1substrings. |
|
|
Split by lines ('\r', '\r\n', '\n') and return a list containing each line as an element. If the argument keepends is False, newline characters are not included; if True, newline characters are preserved. |
|
|
Check whether the string starts with obj; if so, return True, otherwise return False. If beg and end specify values, check within the specified range. |
|
|
Perform lstrip() and rstrip() on string. |
|
|
Swap the case in string. |
|
|
Return a titlecased string, meaning all words start with an uppercase letter and the remaining letters are lowercase (see istitle()). |
|
|
Convert the characters of string according to the table given by str (containing 256 characters), place the characters to be filtered out into the del parameter. |
|
|
Convert lowercase letters in string to uppercase. |
|
|
Return a string of length width, with the original string right-aligned and padded with zeros on the left. |