ansaurus

Question

regex to get value of inside a particular TD in HTML

Answer 1

+7 A:

Please don't use regexes to parse HTML!

Grab a copy of the HTML agility pack and your life will be much simpler, and your application much less brittle.

jvenema 2009-12-09 20:59:03

+1, enough said

just somebody 2009-12-09 21:00:21

Answer 2

A:

Is it arbitrary HTML? Can it have CDATA blocks, comments, external character entities?

If any the above is true, then you should forget about regex for this purpose, and use something like HTML Agility Pack to properly parse it to DOM, and then work with that.

Pavel Minaev 2009-12-09 21:00:17

Answer 3

A:

Use a parser to get the content from the particular TD you care about, then use a regex along the lines of \d of (\d{2})$ which should get the value of xx in the first capture group.

I'm specifically not trying to write a regex that will handle the HTML part of this question; see also the <center> cannot hold.

Hank Gay 2009-12-09 21:02:24

Answer 4

+1 A:

Using regex is not the correct way to do this. As others have pointed out, use an HTML parser. If you have HTML Agility Pack, you can do this:

using System;
using System.Linq;
using System.Text.RegularExpressions;
using HtmlAgilityPack;

class Program
{
    static void Main(string[] args)
    {
        string html = @"<html><body><td class=""blah"" ...........>Some text blah: page 13 of 99<br> more stuff</td></body></html>";
        HtmlDocument doc = new HtmlDocument();
        doc.LoadHtml(html);
        var nodes = doc.DocumentNode.SelectNodes("//td[@class='blah']");
        if (nodes != null)
        {
            var td = nodes.FirstOrDefault();
            if (td != null)
            {
                Match match = Regex.Match(td.InnerText, @"page \d+ of (\d+)");
                if (match.Success)
                {
                    Console.WriteLine(match.Groups[1].Value);
                }
            }
        }
    }
}

Output:

However, it can be done with regex, as long as you accept that it's not going to be a perfect solution. It's fragile, and can easily be tricked, but here it is:

class Program
{
    static void Main(string[] args)
    {
        string s = @"stuff <td class=""blah"" ...........>Some text blah: page 13 of 99<br> more stuff";
        Match match = Regex.Match(s, @"<td[^>]*\sclass=""blah""[^>]*>[^<]*page \d+ of (\d+)<br>");

        if (match.Success)
        {
            Console.WriteLine(match.Groups[1].Value);
        }
    }
}

Output:

Just make sure no-one ever sees you do this.

Mark Byers 2009-12-09 21:03:28

where is the first match?

mrblah 2009-12-09 21:28:55

I'm sorry, I don't understand your question. Can you rephrase it? Which implementation are you referring to: the HTML parsing or the pure regex?

Mark Byers 2009-12-09 21:36:07

could (\d+) be converted into a named match? so I could do match.Groups["somename"].Value ?

mrblah 2009-12-09 21:36:39

Yes. Change `(\d+)` to `(?<somename>\d+)` in either version.

Mark Byers 2009-12-09 21:44:36

ansaurus

tags:

views:

answers:

regex to get value of inside a particular TD in HTML

related questions