ansaurus

Question

Answer 1

+22 A:

You can write a LINQ-based line reader pretty easily using an iterator block:

static IEnumerable<SomeType> ReadFrom(string file) {
    string line;
    using(var reader = File.OpenText(file)) {
        while((line = reader.ReadLine()) != null) {
            SomeType newRecord = /* parse line */
            yield return newRecord;
        }
    }
}

or to make Jon happy:

static IEnumerable<string> ReadFrom(string file) {
    string line;
    using(var reader = File.OpenText(file)) {
        while((line = reader.ReadLine()) != null) {
            yield return line;
        }
    }
}
...
var typedSequence = from line in ReadFrom(path)
                    let record = ParseLine(line)
                    where record.Active // for example
                    select record.Key;

then you have ReadFrom(...) as a lazily evaluated sequence without buffering, perfect for Where etc.

Note that if you use OrderBy or the standard GroupBy, it will have to buffer the data in memory; ifyou need grouping and aggregation, "PushLINQ" has some fancy code to allow you to perform aggregations on the data but discard it (no buffering). Jon's explanation is here.

Marc Gravell 2009-08-13 10:45:02

Bah, separation of concerns - separate out the line reading into a separate iterator, and use normal projection :)

Jon Skeet 2009-08-13 10:48:15

Touché

Marc Gravell 2009-08-13 10:49:07

Much nicer... though still file-specific ;)

Jon Skeet 2009-08-13 10:54:29

aims kick...

Marc Gravell 2009-08-13 11:02:00

I don't think that your examples will compile. "file" is already defined as a string param, so you can't make that declaration in the using block.

fatcat1111 2010-08-07 00:28:51

@fatcat111- fair point - will edit

Marc Gravell 2010-08-07 07:26:49

Answer 2

+9 A:

It's simpler to read a line and check whether or not it's null than to check for EndOfStream all the time.

However, I also have a LineReader class in MiscUtil which makes all of this a lot simpler - basically it exposes a file (or a Func<TextReader> as an IEnumerable<string> which lets you do LINQ stuff over it. So you can do things like:

var query = from file in Directory.GetFiles("*.log")
            from line in new LineReader(file)
            where line.Length > 0
            select new AddOn(line); // or whatever

The heart of LineReader is this implementation of IEnumerable<string>.GetEnumerator:

public IEnumerator<string> GetEnumerator()
{
    using (TextReader reader = dataSource())
    {
        string line;
        while ((line = reader.ReadLine()) != null)
        {
            yield return line;
        }
    }
}

Almost all the rest of the source is just giving flexible ways of setting up dataSource (which is a Func<TextReader>).

Jon Skeet 2009-08-13 10:45:59

Answer 3

+2 A:

NOTE: You need to watch out for the IEnumerable<T> solution, as it will result in the file being open for the duration of processing.

For example, with Marc Gravell's response:

foreach(var record in ReadFrom("myfile.csv")) {
    DoLongProcessOn(record);
}

the file will remain open for the whole of the processing.

David Kemp 2009-08-13 10:50:08

True, but "file open for a long time, but no buffering" is often better than "lots of memory hogged for a long time"

Marc Gravell 2009-08-13 10:53:44

That's true - but basically you've got three choices: load the lot in one go (fails for big files); hold the file open (as you mention); reopen the file regularly (has a number of issues). In many, many cases I believe that streaming and holding the file open is the best solution.

Jon Skeet 2009-08-13 10:55:49

Yes, it's probably a better solution to keep the file open, but you just need to be away of the implication

David Kemp 2009-08-13 11:06:03

Sorry for the name typo there Marc

David Kemp 2009-08-13 11:21:11

It's definitely something to keep in mind as a potentially unexpected side effect, but I also agree with Jon in that it does sound like the best solution.

Mark LeMoine 2010-08-17 17:56:40

Answer 4

A:

Thanks all for your answers! I decided to go with a mixture, mainly focusing on Marc's though as I will only need to read lines from a file. I guess you could argue seperation is needed everywhere, but heh, life is too short!

Regarding the keeping the file open, that isn't going to be an issue in this case, as the code is part of a desktop application.

Lastly I noticed you all used lowercase string. I know in Java there is a difference between capitalised and non capitalised string, but I thought in C# lowercase string was just a reference to capitalised String?

public void Load(AddonCollection<T> collection)
{
    // read from file
    var query =
        from line in LineReader(_LstFilename)
        where line.Length > 0
        select CreateAddon(line);

    // add results to collection
    collection.AddRange(query);
}

protected T CreateAddon(String line)
{
    // create addon
    T addon = new T();
    addon.Load(line, _BaseDir);

    return addon;
}

protected static IEnumerable<String> LineReader(String fileName)
{
    String line;
    using (var file = System.IO.File.OpenText(fileName))
    {
        // read each line, ensuring not null (EOF)
        while ((line = file.ReadLine()) != null)
        {
            // return trimmed line
            yield return line.Trim();
        }
    }
}

Luca Spiller 2009-08-13 16:21:03

Why are you passing the collection into the Load method? At least call it LoadInto if you're going to do that ;)

David Kemp 2009-08-20 09:28:21

ansaurus

tags:

views:

answers:

C# Reading a File Line By Line

related questions