قائمة بمراجع كيانات الأحرف في XML و HTML

في مستندات SGML و HTML و XML ، تتكون البنى المنطقية المعروفة باسم بيانات الأحرف وقيم السمات من تسلسلات من الأحرف، حيث يمكن لكل حرف أن يظهر مباشرةً (ممثلاً نفسه)، أو يمكن تمثيله بسلسلة من الأحرف تُسمى مرجع الأحرف ، والتي يوجد منها نوعان: مرجع أحرف رقمي ومرجع كيان أحرف . تُدرج هذه المقالة مراجع كيانات الأحرف الصالحة في مستندات HTML وXML.

نظرة عامة على مرجع الشخصية

في لغتي HTML وXML، يُشير مرجع الحرف الرقمي إلى حرفٍ ما باستخدام رمزه في مجموعة الأحرف العالمية ( Unicode) ، وذلك بالصيغة التالية: أو حيث يجب أن يكون الحرف صغيرًا في مستندات XML، و هو رمز الحرف بالصيغة الست عشرية، و هو رمز الحرف بالصيغة العشرية . يمكن أن يتكون (أو ) من أي عدد من الأرقام الست عشرية (أو العشرية)، وقد يتضمن أصفارًا بادئة. يمكن أن يجمع رمز الحرف للأرقام الست عشرية بين الأحرف الكبيرة والصغيرة، مع أن الأحرف الكبيرة هي النمط الشائع. تُقيّد معايير XML وHTML رموز الأحرف القابلة للاستخدام بمجموعة من القيم الصالحة، وهي مجموعة فرعية من قيم رموز UCI/Unicode، تستبعد جميع رموز الأحرف غير الحرفية أو البديلة، ومعظم رموز الأحرف المخصصة لعناصر التحكم C0 وC1 (باستثناء فواصل الأسطر وعلامات الجدولة التي تُعامل كمسافات بيضاء).&#xhhhh;&#nnn;xhhhhnnnhhhhnnnhhhh

على النقيض من ذلك، يشير مرجع كيان الأحرف إلى سلسلة من حرف واحد أو أكثر باسم كيان يحتوي على الأحرف المطلوبة كنص بديل . الصيغة هي: &name;حيث nameيمثل اسم الكيان (مع مراعاة حالة الأحرف). عادةً ما تكون الفاصلة المنقوطة مطلوبة في مرجع كيان الأحرف، ما لم يُذكر خلاف ذلك في الجدول أدناه (انظر [ أ ] ). يجب أن يكون الكيان إما مُعرَّفًا مسبقًا (مُدمجًا في لغة الترميز)، أو مُعلنًا عنه في تعريف نوع المستند (DTD) باستخدام [ ب ] .<!ENTITY name "value">

مجموعات الكيانات العامة القياسية للأحرف

XML
تحدد لغة XML خمسة كيانات مُعرَّفة مسبقًا لدعم كل حرف ASCII قابل للطباعة: &amp;، &lt;، ، &gt;، &apos;، و &quot;. الفاصلة المنقوطة في نهاية الكلمة إلزامية في XML (و XHTML ) لهذه الكيانات الخمسة (حتى لو سمحت HTML أو SGML بحذفها لبعضها، وفقًا لتعريف نوع المستند الخاص بها).
مجموعات كيانات ISO
قدّمت لغة SGML مجموعة شاملة من تعريفات الكيانات للأحرف المستخدمة على نطاق واسع في النشر التقني والمرجعي الغربي، وذلك للأحرف اللاتينية واليونانية والسيريلية. كما ساهمت الجمعية الرياضية الأمريكية بكيانات للأحرف الرياضية (انظر [ ج ] ).
مجموعات كيانات HTML
تم بناء الإصدارات المبكرة من لغة HTML باستخدام مجموعات فرعية صغيرة من هذه المجموعات، والتي تتعلق بالأحرف الموجودة في ثلاث مجموعات أحرف غربية ذات 8 بت.
مجموعات كيانات MathML
قام اتحاد شبكة الويب العالمية (W3C) بتطوير مجموعة من تعريفات الكيانات لأحرف MathML .
مجموعات كيانات XML
تولى فريق عمل MathML التابع لاتحاد شبكة الويب العالمية (W3C) مسؤولية صيانة مجموعات الكيانات العامة لمعيار ISO، بالإضافة إلى MathML، وقام بتوثيقها في تعريفات كيانات XML للأحرف . تدعم هذه المجموعة متطلبات XHTML و MathML ، وتُستخدم كمدخل للإصدارات المستقبلية من HTML.
HTML5
تعتمد لغة HTML5 كيانات XML كمراجع أحرف مُسماة ، ولا تُجمّعها في مجموعات. وتستمد أسماء مراجع الأحرف من تعريفات كيانات XML للأحرف. كما تُوفّر مواصفات HTML5 أيضًا روابط بين هذه الأسماء وتسلسلات أحرف Unicode باستخدام JSON .

تم تطوير العديد من مجموعات الكيانات الأخرى لتلبية متطلبات خاصة، وللنصوص الرئيسية والثانوية. ومع ذلك، فقد حل ظهور يونيكود محلها إلى حد كبير.

المعرفات العامة الرسمية لمجموعات فرعية من كيانات HTML DTD

يتم في الواقع تعيين المعرف العام الرسمي الكامل ومعرف النظام لمجموعة فرعية من كيانات DTD (حيث يتم تعريف اسم كيان الأحرف) من أحد الكيانات المسماة الثلاثة التالية المحددة:

مجموعات فرعية من كيانات HTML DTD
اسمإصدارالمعرف العام الرسميمعرّف النظام
HTMLlat1HTML 4"-//W3C//ENTITIES Latin 1//EN//HTML""http://www.w3.org/TR/html4/HTMLlat1.ent"(خياري)
XHTML 1"-//W3C//ENTITIES Latin 1 for XHTML//EN""http://www.w3.org/TR/xhtml1/DTD/xhtml-lat1.ent"
رمز HTMLHTML 4"-//W3C//ENTITIES Symbols//EN//HTML""http://www.w3.org/TR/html4/HTMLsymbol.ent"(خياري)
XHTML 1"-//W3C//ENTITIES Symbols for XHTML//EN""http://www.w3.org/TR/xhtml1/DTD/xhtml-symbol.ent"
HTMLspecialHTML 4"-//W3C//ENTITIES Special//EN//HTML""http://www.w3.org/TR/html4/HTMLspecial.ent"(خياري)
XHTML 1"-//W3C//ENTITIES Special for XHTML//EN""http://www.w3.org/TR/xhtml1/DTD/xhtml-special.ent"
html.dtd [ i ]غير متوفر"http://info.cern.ch/MarkUp/html-spec/html.dtd"
HTML 5 [ ii ]"-//W3C//ENTITIES HTML MathML Set//EN//XML""http://www.w3.org/2003/entities/2007/htmlmathml-f.ent"
  1. ملف تعريف نوع المستند (DTD) الأصلي لـ HTML 1.0، والذي كان متاحًا على الرابط التالي: http://info.cern.ch/MarkUp/html-spec/html.dtd
  2. لا يوجد تعريف نوع المستند (DTD) لـ HTML 5، حيث جميع الكيانات مُعرّفة مسبقًا؛ من المستحيل التحقق بدقة من صحة المخطط المطلوب لـ (X)HTML 5 في XML، دون تعريف XSD مخصص (على الأقل لسمات "data-*" المخصصة). بدلًا من اشتراط دعم تعريف نوع المستند (DTD) (وما يرتبط به من مخاوف أمنية ) ، فإن أفضل طريقة لتبادل HTML5 مع XHTML بشكل آمن هي تحويل جميع مراجع الكيانات إلى مراجع نصية عادية، أو مراجع رقمية، أو (حيثما ينطبق) الكيانات الخمسة القياسية لـ XML 1.0. مع ذلك:
    • تُستخدم مجموعة كيانات HTML 5 أيضًا في MathML 3، ولهذا الغرض، يتم تعيين مجموعة المعرفات لمجموعة كيانات DTD الخاصة بها . [ 1 ]PUBLIC "-//W3C//ENTITIES HTML MathML Set//EN//XML" "http://www.w3.org/2003/entities/2007/htmlmathml-f.ent"
    • تشجع مواصفات WHATWG المتصفحات على ربط المعرفات العامة الرسمية لـ MathML 2 أو XHTML 1.x (عند استخدامها في XML) بمعرف بيانات URI يحتوي على مجموعة كيانات HTML5، وإعطاء هذه الأولوية على معرف النظام المقدم، وذلك من أجل "التعامل مع الكيانات بطريقة قابلة للتشغيل البيني دون الحاجة إلى أي وصول إلى الشبكة". [ 2 ]

المعرفات العامة الرسمية لمجموعات فرعية من كيانات ISO القديمة

مجموعات كيانات ISO هي مجموعات فرعية قديمة (موثقة) من الأحرف، والتي تم إعطاؤها أسماء كيانات أحرف SGML في ISO 8879 وISO 9573، والتي تم استخدامها في الترميزات القديمة قبل التوحيد ضمن ISO 10646. معرفاتها العامة الرسمية الكاملة هي كما يلي:

مجموعات فرعية من كيانات ISO
اسمالمعرف (المعرفات) العامة الرسمية
ISOamsa
  • "ISO 8879:1986//ENTITIES Added Math Symbols: Arrow Relations//EN"[ i ] [ 3 ]
  • "ISO 9573-13:1991//ENTITIES Added Math Symbols: Arrow Relations//EN"[ 4 ]
ISOamsb
  • "ISO 8879:1986//ENTITIES Added Math Symbols: Binary Operators//EN"[ i ] [ 3 ]
  • "ISO 9573-13:1991//ENTITIES Added Math Symbols: Binary Operators//EN"[ 4 ]
ISOamsc
  • "ISO 8879:1986//ENTITIES Added Math Symbols: Delimiters//EN"[ i ] [ 3 ]
  • "ISO 9573-13:1991//ENTITIES Added Math Symbols: Delimiters//EN"[ 4 ]
ISOamsn
  • "ISO 8879:1986//ENTITIES Added Math Symbols: Negated Relations//EN"[ i ] [ 3 ]
  • "ISO 9573-13:1991//ENTITIES Added Math Symbols: Negated Relations//EN"[ 4 ]
ISOamso
  • "ISO 8879:1986//ENTITIES Added Math Symbols: Ordinary//EN"[ i ] [ 3 ]
  • "ISO 9573-13:1991//ENTITIES Added Math Symbols: Ordinary//EN"[ 4 ]
ISOamsr
  • "ISO 8879:1986//ENTITIES Added Math Symbols: Relations//EN"[ i ] [ 3 ]
  • "ISO 9573-13:1991//ENTITIES Added Math Symbols: Relations//EN"[ 4 ]
ISObox"ISO 8879:1986//ENTITIES Box and Line Drawing//EN"[ i ] [ 3 ]
إيزوكيم"ISO 9573-13:1991//ENTITIES Chemistry//EN"[ 4 ]
ISOcyr1"ISO 8879:1986//ENTITIES Russian Cyrillic//EN"[ i ] [ 3 ]
ISOcyr2"ISO 8879:1986//ENTITIES Non-Russian Cyrillic//EN"[ i ] [ 3 ]
ISOdia"ISO 8879:1986//ENTITIES Diacritical Marks//EN"[ i ] [ 3 ]
ISOgrk1"ISO 8879:1986//ENTITIES Greek Letters//EN"[ i ] [ 3 ]
ISOgrk2"ISO 8879:1986//ENTITIES Monotoniko Greek//EN"[ i ] [ 3 ]
ISOgrk3
  • "ISO 8879:1986//ENTITIES Greek Symbols//EN"[ i ] [ 3 ]
  • "ISO 9573-13:1991//ENTITIES Greek Symbols//EN"[ 4 ]
ISOgrk4
  • "ISO 8879:1986//ENTITIES Alternative Greek Symbols//EN"[ i ] [ 3 ]
  • "ISO 9573-13:1991//ENTITIES Alternative Greek Symbols//EN"[ 4 ]
ISOlat1"ISO 8879:1986//ENTITIES Added Latin 1//EN"[ i ] [ ii ] [ 3 ]
ISOlat2"ISO 8879:1986//ENTITIES Added Latin 2//EN"[ i ] [ 3 ]
ISOmfrk"ISO 9573-13:1991//ENTITIES Math Alphabets: Fraktur//EN"[ 4 ]
ISOmopf"ISO 9573-13:1991//ENTITIES Math Alphabets: Open Face//EN"[ 4 ]
ISOmscr"ISO 9573-13:1991//ENTITIES Math Alphabets: Script//EN"[ 4 ]
ISOnum"ISO 8879:1986//ENTITIES Numeric and Special Graphic//EN"[ i ] [ 3 ]
ISOpub"ISO 8879:1986//ENTITIES Publishing//EN"[ i ] [ 3 ]
ISOtech
  • "ISO 8879:1986//ENTITIES General Technical//EN"[ i ] [ 3 ]
  • "ISO 9573-13:1991//ENTITIES General Technical//EN"[ 4 ]
  1. 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 يُعتبرالإصدار الذي يبدأ بـISO 8879-1986//بدلاً من الإصدار القديم مُهملاً. [ 3 ]ISO 8879:1986//
  2. يتم استخدامنسخة مع الملحق أحيانًا بشكل خاطئ لمجموعة كيانات HTMLlat1 الأكبر ، أي بدلاً من [ 3 ] (انظر أعلاه ).//HTML"-//W3C//ENTITIES Latin 1//EN//HTML"

قائمة بمراجع كيانات الأحرف في لغة HTML

يُعرّف HTML5 العديد من الكيانات المُسماة، والتي تُستخدم كأسماء بديلة مُختصرة لبعض أحرف Unicode. [ 5 ] لا تسمح مواصفات HTML5 للمستخدمين بتعريف كيانات إضافية، حيث لم تعد تقبل أي تعريف نوع المستند (DTD) للإشارة إليه أو توسيعه داخل مستندات HTML (لا يزال هذا مطلوبًا في XHTML، الذي يعتمد على قواعد تحليل XML أكثر صرامة ولكنه يسمح بالإشارة إلى تعريف DTD أو تعريفه في رأس المستند، لأن XML لا يُعرّف مُسبقًا مُعظم كيانات HTML).

في الجدول أدناه، يشير عمود "القياسي" إلى الإصدار الأول من تعريف نوع المستند (DTD) الخاص بـ HTML الذي يُعرّف مرجع كيان الحرف، ويشير إلى الأحرف المُعرّفة مسبقًا في XML دون الحاجة إلى أي تعريف نوع مستند. لاستخدام أحد مراجع كيانات الأحرف هذه في مستند HTML أو XML، أدخل علامة العطف (&) متبوعة باسم الكيان ، ثم فاصلة منقوطة (إلزامية في XML، وموصى بها بشدة في HTML لجميع الكيانات، حتى لو كان HTML يسمح بحذف الفاصلة المنقوطة من بعض الكيانات فقط، كما هو موضح أدناه بـ [ a ] ). على سبيل المثال، أدخل &copy;للرمز © لحقوق النشر .

لا توجد كيانات أحرف مُعرَّفة مسبقًا في لغة HTML للأحرف أو تسلسلات معظم النصوص المُشفَّرة في نظام ترميز Unicode (باستثناء مجموعة فرعية شائعة من المسافات البيضاء، وعلامات الترقيم، والرموز الرياضية أو التقنية، ورموز العملات، وبعض الرموز العبرية المستخدمة في التدوينات الرياضية، وأكثر الأحرف شيوعًا في اللاتينية أو اليونانية أو السيريلية). تجدر الإشارة أيضًا إلى أنه لا يتم تمثيل جميع عناصر التحكم ثنائية الاتجاه المُعرَّفة في نظام ترميز Unicode ككيانات أحرف قياسية في لغة HTML (ولا حتى في HTML5، الذي يُعرِّف عناصر وسمات اتجاهية أكثر عمومية لهذا الغرض). ومن الجدير بالذكر أنه لا توجد كيانات أحرف مُعرَّفة مسبقًا في لغة HTML لعناصر التحكم التي أُضيفت في نظام ترميز Unicode وتم تعريفها رسميًا في الإصدار الثاني من خوارزمية Unicode ثنائية الاتجاه.

معظم الكيانات مُعرّفة مسبقًا في XML وHTML للإشارة إلى حرف واحد فقط في نظام ترميز Unicode (UCS)، ولكن لا توجد كيانات مُعرّفة مسبقًا للأحرف المركبة المنفردة، أو مُحدِّدات التباين، أو الأحرف المُخصصة للاستخدام الخاص؛ ومع ذلك، تتضمن القائمة بعض الكيانات المُعرّفة مسبقًا لتسلسلات الأحرف المكونة من حرفين والتي تحتوي على بعضها. منذ HTML 5.0 (وMathML 3.0 التي تشترك في نفس مجموعة الكيانات)، يتم ترميز جميع الكيانات باستخدام صيغ توحيد Unicode C وKC (لم يكن هذا هو الحال مع الإصدارات الأقدم من HTML وMathML، لذلك تم تعديل الكيانات القديمة التي تم تعريفها في البداية باستخدام أحرف مُخصصة للاستخدام الخاص، أو صيغ توافق CJK، أو صيغ غير متوافقة مع معيار NFC [ 6 ] ).

ومع ذلك، فإن جميع الأحرف والتسلسلات الصالحة في نظام الإحداثيات الموحد، بما في ذلك جميع عناصر التحكم ثنائية الاتجاه أو تعيينات الاستخدام الخاص (ولكن باستثناء عناصر التحكم C0 و C1 غير البيضاء، والأحرف غير، والبدائل) قابلة للاستخدام وصالحة أيضًا في HTML و XML و XHTML و MathML، سواء في قيم النص العادي للسمات أو في عناصر النص (عن طريق ترميزها مباشرة كنص عادي، أو باستخدام مراجع الأحرف الرقمية عند الحاجة).

Notes

  1. 123456789101112131415161718192021222324252627282930313233343536373839404142434445464748495051525354555657585960616263646566676869707172737475767778798081828384858687888990919293949596979899100101102103104105106107108The trailing semicolon may be omitted for this named entity.
  2. 12 DTD: see § Formal public identifiers for HTML DTD entities subsets
  3. 12Old ISO subset: see § Formal public identifiers for old ISO entities subsets
  4. Description: the standard ISO 10646 and Unicode character name is displayed first for each character, with non-standard but legacy synonyms shown in italics between parentheses after an equal sign.
  5. 123The leading space before combining characters used in old DTDs for MathML2.0 was removed in MathML 3.0 and HTML 5.0.
  6. &quot; was omitted from the HTML 3.2 specification, but was restored as of HTML 4.0. In practice, most web browsers displaying HTML 3.2 pages render it as if it had been included in the spec.
  7. 1234spaces: a blue background is used to display each space's width.
  8. &copy;: U+00A9 'copyright symbol' is not the same as U+24B8 'circled Latin capital letter C', although the same glyph could be used do depict both characters. See also U+24D2 'Latin small letter c'.
  9. &reg;: U+00AE 'registered sign' is not the same as U+24C7 'circled Latin capital letter R', although the same glyph could be used do depict both characters.
  10. &angst;: The use of U+212B 'Angstrom sign', which was encoded due to round-trip mapping compatibility with an East-Asian character encoding, is discouraged, and the preferred representation is U+00C5 'capital letter A with ring above', which has the same glyph.
  11. 12&IJlig; and &ijlig;: The use of U+0132 'IJ ligature' or U+0133 'ij ligature', which were encoded for usage in Dutch and for compatibility for ISO/IEC 6937 and Code page 1102 (which only includes the lowercase ij, also part of the Dutch version of ISO 646 National Replacement Character Set), is discouraged, and the preferred representation is simply 'IJ' or 'ij' (as two separate letters).
  12. 12&lmidot;: The use of U+013F 'Latin small letter l with middle dot' or U+0140 'Latin capital letter L with middle dot', which were encoded for usage in Catalan and for compatibility for ISO/IEC 6937, is discouraged, and the preferred representation is 'L' or 'l', followed by U+00B7.
  13. &napos;: The use of U+0149 'n preceded by apostrophe', which was encoded for usage in Afrikaans and for compatibility for ISO/IEC 6937, has been deprecated by Unicode (since Unicode 5.2), and the preferred representation is ʼn (U+02BC followed by n). (Unicode.org – Proposal for Additional Deprecated Characters).
  14. 12ligature: this is a standard misnomer as this is a separate character in some languages.
  15. 12345678910111213141516171819202122232425262728293031323334353637383940414243444546474849Greek letters: the ISOgrk1 set includes a set of entity names for the entire Greek alphabet (without diacritics),[7] while the ISOgrk3 set includes a different set of entity names for the subset of the Greek letters used contrastively with Latin letters in mathematical notation.[8] The HTML HTMLsymbol set includes an expanded version of the ISOgrk3 set, not the ISOgrk1 set.
  16. &ohm;: The use of U+2126 'ohm sign', is discouraged, and the preferred representation is U+03A9 'Greek capital letter Omega', which has the same glyph.
  17. 1234&NegativeMediumSpace;, &NegativeThickSpace;, &NegativeThinSpace;, &NegativeVeryThinSpace;: these are names used in the Wolfram Language for Private Use Area characters with negative advance widths;[9][10][11][12] HTML5 approximates them with the zero-width space.
  18. 12345black: here it seems to mean filled as opposed to hollow.
  19. 12ISO proposed: these characters have been standardized in ISO 10646 after the release of HTML 4.0.
  20. 1234&image;, &map;: these two entity names were defined differently, as file-type icons, in the abandoned specification for HTML version 3.0.[13][14]
  21. &copysr;: U+2117 'sound recording copyright' is not the same as U+24C5 'circled Latin capital letter P', although the same glyph could be used do depict both characters.
  22. &alefsym;: U+2135 'alef symbol' is not the same as U+05D0 'Hebrew letter alef' (which, unlike the mathematical symbol, has strong right-to-left bidirectional text behaviour), although the same glyph could be used to depict both characters.
  23. &beth;: U+2136 'bet symbol' is not the same as U+05D1 'Hebrew letter bet' (which, unlike the mathematical symbol, has strong right-to-left bidirectional text behaviour), although the same glyph could be used to depict both characters.
  24. &gimel;: U+2137 'gimel symbol' is not the same as U+05D2 'Hebrew letter gimel' (which, unlike the mathematical symbol, has strong right-to-left bidirectional text behaviour), although the same glyph could be used to depict both characters.
  25. &daleth;: U+2138 'dalet symbol' is not the same as U+05D3 'Hebrew letter dalet' (which, unlike the mathematical symbol, has strong right-to-left bidirectional text behaviour), although the same glyph could be used to depict both characters.
  26. &lArr;: ISO 10646 does not say that 'leftwards double arrow' is the same as the 'is implied by' arrow, but also does not have any other character for that function, so lArr can be used for 'is implied by' as ISOtech suggests.
  27. &rArr;: ISO 10646 does not say that 'rightwards double arrow' is the same as the 'implies' arrow, but also does not have any other character with this function, so rArr can be used for 'implies' as ISOtech suggests.
  28. &prod;: U+220F 'n-ary product' is not the same character as U+03A0 'Greek capital letter Pi' though the same glyph might be used for both.
  29. &sum;: U+2211 'n-ary summation' is not the same character as U+03A3 'Greek capital letter Sigma' though the same glyph might be used for both.
  30. &sim;: U+223C 'tilde operator' is not the same character as U+007E 'tilde', although the same glyph might be used to represent both.
  31. &nsup;: U+2285 'not a superset of' is in the 'ISOamsn' subset, but is not covered by the Symbol font encoding, and is not listed in the HTML 4.0 entities list on the documentation, where it was erroneously omitted; it should be included for symmetry and analogy with other entities.
  32. &perp;: Unicode only defines U+22A5 as the "up tack", and the Unicode symbol for "perpendicular" is U+27C2: the two symbols look similar, but are separate in Unicode. However, HTML uses U+22A5 as its "perpendicular" symbol: this is a discrepancy between HTML and Unicode. As well, the U+22A4 character (the "down tack" symbol) rendered in a browser such as Firefox 3.6 can match the font of either "up tack" or "perpendicular", but not both, depending on whether a fixed-width or a proportional font is used. When viewed in Firefox 3.6, the symbols rendered in the order U+22A5, U+22A4, U+27C2 in a proportional font: "⊥ ⊤ ⟂" and a fixed width one: ⊥ ⊤ ⟂, shows that the "down tack" has a similar look to U+22A5 (HTML's "perpendicular") in the first case but matches U+27C2 in the second. This exemplifies the difficulties of the semiotics involved in interpreting glyphs, symbols and characters generally.
  33. &sdot;: U+22C5 'dot operator' is not the same character as U+00B7 'middle dot'.
  34. &Ll;: U+22D8 'very much less-than' is missing in the HTML 5.2 list of entities, where it was omitted.
  35. &lang;: U+27E8 'mathematical left angle bracket' is not the same character as U+003C 'less than', U+2039 'single left-pointing angle quotation mark', or U+3008 'left angle bracket'. In HTML 5.0, lang was remapped to this code, as U+2329 'left-pointing angle bracket' has been marked deprecated in Unicode (since version 5.2) (Unicode.org – Proposal for Additional Deprecated Characters).
  36. &rang;: U+27E9 'mathematical right angle bracket' is not the same character as U+003E 'greater than', U+203A 'single right-pointing angle quotation mark', or U+3009 'right angle bracket'. In HTML 5.0, rang had been remapped to this code, as U+232A 'right-pointing angle bracket' has been marked deprecated in Unicode (since version 5.2) (Unicode.org – Proposal for Additional Deprecated Characters).

Entities representing special characters in XHTML

The XHTMLDTDs explicitly declare 253 entities (including the 5 predefined entities of XML 1.0) whose expansion is a single character, which can therefore be informally referred to as "character entities". These (with the exception of the &apos; entity) have the same names and represent the same characters as the 252 character entities in HTML 4.0. Also, by virtue of being XML, XHTML documents may reference the predefined &apos; entity, which is not one of the 252 character entities in HTML 4.0. Additional entities of any size may be defined on a per-document basis. However, the usability of entity references in XHTML is affected by how the document is being processed:

  • Legacy abbreviated character entities (without the final colon) inherited from HTML 2.0 (and still supported in HTML 5.0) are not supported in XML 1.0 and XHTML; the trailing semicolon must be present in all entity references used in XML and XHTML documents.
  • If the XHTML document is read by a conforming HTML 4.0 processor, then only the 252 HTML 4.0 character entities may safely be used. The use of &apos; or custom entity references may not be supported and may produce unpredictable results (it is recommended to use the numerical character reference &#39; instead).
  • If the document is read by an XML parser that does not or cannot read external entities, then only the five built-in XML character entities can safely be used, although other entities may be used if they are declared in the internal DTD subset. However, modern XML parsers recognize and implement a builtin cache for SGML references to DTDs used by all standard versions of HTML, XHTML, SVG and MathML, without needing to parse and process the external DTD via their URL and without needing to process entities defined in an internal DTD subset of the document.
  • If the document is read by an XML parser that does read external entities and does not implement a builtin cache for well-known DTDs, then the five built-in XML character entities (and numeric character references) can safely be used. The other 248 HTML character entities can be used as long as the XHTML DTD is accessible to the parser at the time the document is read. Other entities may also be used if they are declared in the internal DTD subset and the XML processor can parse internal DTD subsets.
  • HTML 5.0 parsers cannot process XHTML documents, and it's impossible to define a fully validating DTD for HTML5 documents encoded with the XHTML syntax (notably it's impossible to validate all attributes names, notably "data-*" attributes); as well it's still impossible to fully validate (with W3C standard schemas for XML, such as XSD or relax NG) HTML5 documents represented in the XHTML syntax, and for now a custom validator specific to HTML 5.0 is required.

Because of the special &apos; case mentioned above, only &quot;, &amp;, &lt;, and &gt; will work in all XHTML processing situations.

See also

References

  1. "htmlmathml-f entity set". W3C. 2011.
  2. "14.2 Parsing XML documents". HTML Standard. WHATWG. Retrieved 13 July 2024.
  3. 123456789101112131415161718192021"sgml-iso-entities-8879.1986/catalog". Debian. 2013.
  4. 12345678910111213"sgml-iso-entities-9573-13.1991/catalog". Debian. 2013.
  5. "HTML5 Named Character Reference List".
  6. "XML Entity Definitions for Characters (3rd Edition) - § C Differences between these entities and earlier W3C DTDs".
  7. Organization for the Advancement of Structured Information Standards (OASIS) (2002). "ISO Greek Letters Entities V0.3". Debian.
  8. Organization for the Advancement of Structured Information Standards (OASIS) (2002). "ISO Greek Symbols Entities V0.3". Debian.
  9. Wolfram. "\[NegativeThickSpace]". Wolfram Language Documentation.
  10. Wolfram. "\[NegativeMediumSpace]". Wolfram Language Documentation.
  11. وولفرام . "\ [ NegativeThinSpace ] " . توثيق لغة وولفرام .
  12. وولفرام . "\ [ مسافة ضئيلة للغاية سلبية ] " . توثيق لغة وولفرام .
  13. هانا، مايكل ج. (7 ديسمبر 1995). "أيقونات HTML: أسماء كيانات أيقونات HTML المقترحة" . مؤرشف من الأصل في 2 فبراير 2015.
  14. "أيقونات ISO/WWW القياسية مقدمة من بيرت بوس وكيفن هيوز" . W3C .

للمزيد من القراءة