tags:

views:

58

answers:

4

hello

I want to count how many characters a certain string has in PHP, but I cant get it to work.

I did a var_dump and this is what I got:

string(23) "Children’s Center"

It seems like ' gets translated to ’s

This makes it impossible to get the actual character count. I tried using html_entity_decode but it did not help.

Anyone?

EDIT: My function looks like this now:

function make_shorter($string, $maxlength)
{
 $string = stripslashes($string);   
 $string = html_entity_decode($string);
 if ( mb_strlen(utf8_decode($string)) > $maxlength)
  return substr($string,0,$maxlength).'...';
 else
  return $string;
}

I can't change how the data gets in to the system. I can only modify it's output.

(It's a wordpress site)

A: 
echo strlen($myString); 

or

echo strlen(utf8_decode($myString)); 

?

Fosco
A: 

Try strlen(), It will solve

Centurion
that didnt work.
vick
+1  A: 

You're using a right single quotation mark, instead of the ' character. Well... stop that!

$string = "Children's Center";

echo strlen($string); // 17

Google this to see the difference: ’ '

nush
+2  A: 

It seems like ' gets translated to ’s

That's the HTML character reference for the single right smart-quote, ’, which some people also use to represent an apostrophe (especially those typing in MS Word). At some prior point in your processing, something has applied htmlentities() to your data.

HTML-escaping should only be carried out at the output stage, when inserting text into an HTML page. So look back and see if you've got a function doing something stupid like calling htmlentities() over every entry in the $_POST/$_GET/$_REQUEST array. This is a common but completely bogus way to try to prevent XSS attacks. If you see it, you are at some point going to need to take it out and go through your templates adding proper HTML-escaping around every time you drop a variable into HTML.

Either way, use htmlspecialchars() instead of htmlentities() and it'll leave the non-ASCII characters like ’ alone.

html_entity_decode() definitely should undo the above encoding, leaving you with a raw string. ’ might still count as either one or two bytes, though, depending on what encoding you decode into. If you want to count characters properly and you have UTF-8 strings, you will need mb_strlen().

bobince
This seems to be the best answer, but I've been looking at it and, for the life of me, can't get html_entity_decode() to decode the ᾿ properly, even with ENT_QUOTES and UTF-8 as the decoding. Have you had more success?
cincodenada
Works for me with UTF-8: http://codepad.org/RADQmYUj.
bobince
i tried everything, nothing seems to work, check out my code above.
vick
Huh, I must have been doing strange things with mine then.
cincodenada
Pass `utf-8` to the `mb_` functions. Use `mb_substr` instead of `substr`. Working: http://ideone.com/iRquj. (And again, `stripslashes()` is another sign that the input is bogus. If it is using magic quotes, that's another disaster.)
bobince