tags:

views:

158

answers:

5

I have a text document that contains a list of numbers and I want to convert it to a list. Right now I can only get the entire list in the 0th entry of the list, but I want each number to be an element of a list. Does anyone know of an easy way to do this in Python?

1000
2000
3000
4000

to

['1000','2000','3000','4000']
+1  A: 
>>> open("myfile.txt").readlines()
>>> lines = open("myfile.txt").readlines()
>>> lines
['1000\n', '2000\n', '3000\n', '4000\n']
>>> clean_lines = [x.strip() for x in lines]
>>> clean_lines
['1000', '2000', '3000', '4000']

Or, if you have a string already, use str.split:

>>> myfile
'1000\n2000\n3000\n4000\n'
>>> myfile.splitlines()
['1000', '2000', '3000', '4000', '']

You can remove the empty element with a list comprehension (or just a regular for loop)

>>> [x for x in myfile.splitlines() if x != ""]
['1000', '2000', '3000', '4000']
dbr
use `s.splitlines()`, not `s.split("\n")`
Devin Jeanpierre
+6  A: 

To convert a Python string into a list use the str.split method:

>>> '1000 2000 3000 4000'.split()
['1000', '2000', '3000', '4000']

split has some options: look them up for advanced uses.

You can also read the file into a list with the readlines() method of a file object - it returns a list of lines. For example, to get a list of integers from that file, you can do:

lst = map(int, open('filename.txt').readlines())

P.S: See some other methods for doing the same in the comments. Some of those methods are nicer (more Pythonic) than mine

Eli Bendersky
you're using `str.split`, not `string.split`, the latter is obsolete anyway.
SilentGhost
@SilentGhost: typo fixed, thanks for noticing
Eli Bendersky
File objects are iterable, so there is seldom a reason to use the `readlines` method. For example, you could get the same results as your last snippet with `map(int, open('filename.txt'))`.
Mike Graham
List comprehension is usually faster than `map` built-in function plus if the file is big enough iteration over file will be faster and more memory efficient than `readlines`:`lst = [int(line) for line in open('filename.txt')]`
Ruslan Spivak
`lst = [int(line) for line in open('filename.txt')]` is much better (and more efficient) idiomatic Python for your second example. *chuckle* I should read the other comments before I post one.
Omnifarious
A: 

You might need to strip newlines.

# list of strings
[number for number in open("file.txt")]

# list of integers
[int(number) for number in open("file.txt")]
Nikola Smiljanić
And, if the OP were to need to strip newlines, how would that code look?
Omnifarious
You can also just use the `list` builtin instead of a list comprehension - `list(open("myfile.txt"))` -> `['1000\n', '2000\n', '3000\n', '4000\n']`
dbr
@dbr: `.readlines()` might be a better option.
SilentGhost
@SilentGhost True, it makes it clearer what is happening, but `list(open("myfile"))` is the same as `open("myfile").readlines()` - by default iterating over a file with use the `readlines` method
dbr
+1  A: 
    $ cat > t.txt
    1
    2
    3
    4
    ^D
    $ python
    Python 2.6.1 (r261:67515, Jul  7 2009, 23:51:51) 
    [GCC 4.2.1 (Apple Inc. build 5646)] on darwin
    Type "help", "copyright", "credits" or "license" for more information.
    >>> l = [l.strip() for l in open('t.txt')]
    >>> l
    ['1', '2', '3', '4']
    >>> 
pulegium
you don't need to do `.readlines()`!
SilentGhost
true. removed now, thanks for pointing that out!
pulegium
A: 
   with open('file.txt', 'rb') as f:
       data = f.read()
   lines = [s.strip() for s in data.split('\n') if s]
rlotun